0% found this document useful (0 votes)
14 views10 pages

Multimodal Rag

The document introduces FinTMMBench, a benchmark for evaluating temporal-aware multi-modal Retrieval-Augmented Generation (RAG) systems in finance, utilizing diverse data types from NASDAQ 100 companies. It highlights the need for temporal awareness in financial analysis and proposes a novel TMMHybridRAG method that integrates various data modalities and temporal information. The benchmark comprises 5,676 questions and aims to address existing gaps in current financial RAG evaluations by providing a comprehensive assessment framework.

Uploaded by

kenlee.reb
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views10 pages

Multimodal Rag

The document introduces FinTMMBench, a benchmark for evaluating temporal-aware multi-modal Retrieval-Augmented Generation (RAG) systems in finance, utilizing diverse data types from NASDAQ 100 companies. It highlights the need for temporal awareness in financial analysis and proposes a novel TMMHybridRAG method that integrates various data modalities and temporal information. The benchmark comprises 5,676 questions and aims to address existing gaps in current financial RAG evaluations by providing a comprehensive assessment framework.

Uploaded by

kenlee.reb
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Towards Temporal-Aware Multi-Modal Retrieval Augmented

Generation in Finance
Fengbin Zhu∗ Junfeng Li∗ Liangming Pan†
National University of Singapore National University of Singapore Peking University
Singapore Singapore China
zhfengbin@[Link] lijunfeng@[Link] peterpan10211020@[Link]

Wenjie Wang Fuli Feng Chao Wang


University of Science and Technology University of Science and Technology 6Estates Pte Ltd
of China of China Singapore
arXiv:2503.05185v2 [[Link]] 3 Aug 2025

China China wangchao@[Link]


wenjiewang96@[Link] fulifeng93@[Link]

Huanbo Luan Tat-Seng Chua


6Estates Pte Ltd National University of Singapore
Singapore Singapore
luanhuanbo@[Link] chuats@[Link]

Abstract CCS Concepts


Finance decision-making often relies on in-depth data analysis • Information systems → Information retrieval.
across various data sources, including financial tables, news arti-
cles, stock prices, etc. In this work, we introduce FinTMMBench, Keywords
the first comprehensive benchmark for evaluating temporal-aware
Retrieval-Augmented Generation, Temporal-aware Retrieval, Multi-
multi-modal Retrieval-Augmented Generation (RAG) systems in
modal Retrieval, Multi-modal LLM
finance. Built from heterologous data of NASDAQ 100 companies,
FinTMMBench offers three significant advantages. 1) Multi-modal ACM Reference Format:
Corpus: It encompasses a hybrid of financial tables, news articles, Fengbin Zhu, Junfeng Li, Liangming Pan, Wenjie Wang, Fuli Feng, Chao
daily stock prices, and visual technical charts as the corpus. 2) Wang, Huanbo Luan, and Tat-Seng Chua. 2018. Towards Temporal-Aware
Temporal-aware Questions: Each question requires the retrieval and Multi-Modal Retrieval Augmented Generation in Finance. In Proceedings of
interpretation of its relevant data over a specific time period, in- Make sure to enter the correct conference title from your rights confirmation
cluding daily, weekly, monthly, quarterly, and annual periods. 3) emai (Conference acronym ’XX). ACM, New York, NY, USA, 10 pages. https:
Diverse Financial Analysis Tasks: The questions involve 10 different //[Link]/[Link]
financial analysis tasks designed by domain experts, including infor-
mation extraction, trend analysis, sentiment analysis and event de-
tection, etc. We further propose a novel TMMHybridRAG method, 1 Introduction
which first leverages a multi-modal LLM to convert data from other Financial analysis is fundamental to modern finance, supporting ap-
modalities (e.g., tabular, visual and time-series data) into textual plications such as equity investment [9], portfolio optimization [20],
format and then incorporates temporal information in each node and risk management [21]. Effective decision-making in these areas
when constructing graphs and dense indexes. Its effectiveness has requires synthesizing up-to-date information from diverse modali-
been validated in extensive experiments, but notable gaps remain, ties, including structured tables, unstructured text, time-series data,
highlighting the challenges presented by our FinTMMBench. The and visual charts, as illustrated in Figure 1 (a).
benchmark and source code will be made publicly available1 . Recently, Retrieval-Augmented Generation (RAG) systems have
∗ Equal
been increasingly explored in financial analysis [17, 26]. Current
Contribution
† Corresponding Author
financial benchmarks for evaluating RAG systems include FinTex-
1 [Link] tQA [3], AlphaFin [17], OmniEval [26], and FinanceBench [13].
However, these datasets offer limited data modalities, potentially
Permission to make digital or hard copies of all or part of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed harming the validity of evaluation. Specifically, FinTextQA and
for profit or commercial advantage and that copies bear this notice and the full citation OmniEval are restricted to textual data, whereas AlphaFin covers
on the first page. Copyrights for components of this work owned by others than the textual and time-series data, and FinanceBench combines textual
author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or
republish, to post on servers or to redistribute to lists, requires prior specific permission and visual data. In addition, they often fail to adequately incorpo-
and/or a fee. Request permissions from permissions@[Link]. rate temporal information in their task design, which is critical for
Conference acronym ’XX, Woodstock, NY assessing whether RAG systems can accurately retrieve and pro-
© 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM.
ACM ISBN 978-1-4503-XXXX-X/18/06 cess financial data within specific time periods. Although AlphaFin
[Link] introduces some temporal questions, they are solely centered on
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.

Buy, Sell or Hold? employs an LLM to generate descriptions for each entity and re-
lation. For non-textual data, TMMHybridRAG regards each table,
Fundamental Analysis Technical Analysis daily stock price record, and chart as a distinct entity and utilizes
• Debt-to-Equity Ratio • Relative Strength Index
• P/E Ratio • Moving Average an advanced multi-modal LLM to generate a textual summary for
• Market Sentiment • Trend
• ... • ... each, which serves as the entity’s description. Further, TMMHy-
bridRAG integrates temporal information into every entity and
Financial Financial Stock Technical relation as the properties to construct dense vectors and graphs.
Tables News Prices Charts
During prediction, given a question, all retrieved entities and re-
(a) Data-driven Equity Investment
lations from both dense vectors and graphs, along with their raw
Apple Inc. Income Statement Table Question: data, are fed into a multi-modal LLM to infer the answer. Extensive
Fiscal Period Dec 2022 Apr 2023 Jul 2023 What is the Price-to-Earnings
(P/E) ratio of Apple on Dec. experiments show that our TMMHybridRAG method significantly
Total Revenue 117,154.00 94,836.00 81,797.00
30, 2022, given 1,000,000 outperforms all compared methods across all evaluation metrics.
Gross Profit 50,332.00 41,976.00 36,413.00
shares?
Net Income 29,998.00 24,160.00 19,881.00 However, its F1 score remains relatively low at 31.41, highlight-
Operating Income 36,016.00 28,318.00 22,998.00 Supporting Evidence:
The close price of Apple on ing the substantial challenges presented in FinTMMBench and
Apple Inc. Stock Price
132
Dec. 30, 2022, is 129.93 USD. underscoring the need for more advanced RAG methods.
130 Date Close Price The net income of Apple in
128 2022Q4 is 29,998.00 USD. In summary, our major contributions are threefold: 1) To the best
126 Dec 29 2022 129.61
124
Dec 30 2022 129.93 Answer: 4331.29 of our knowledge, we are the first to investigate temporal-aware
122
Jan 03 2023 125.07 Explanation: P/E Ratio = multi-modal RAG in the financial domain, addressing a critical real-
23

23

23
2

2
02

02

02

02

20

20

20

Stock Price / Earning per


72

82

92

02

5
c2

c2

c2

c3

n0

n0

n0

... ... world need in financial analysis. 2) We introduce a new benchmark,


De

De

De

De

Ja

Ja

Ja

Close Price share ...


(b) An Example from the FinTMMBench FinTMMBench, specially designed to evaluate temporal-aware
Figure 1: (a) Illustration of financial analysis for decision- multi-modal RAG systems in finance. FinTMMBench comprises
making. (b) An example from FinTMMBench. 5,676 temporal-aware questions that require information from four
distinct modalities, i.e. financial tables, news articles, daily stock
prices, and visual technical charts, to be answered. 3) To tackle
time-series data. Their narrow focus restricts their ability to com- the challenges in FinTMMBench, we propose TMMHybridRAG,
prehensively evaluate RAG systems in handling temporal-aware a novel temporal-aware multi-modal RAG method that integrates
queries over heterogeneous data across different modalities. dense and graph retrieval techniques. Experiments demonstrate
To address these gaps,we introduce FinTMMBench, a financial that our TMMHybridRAG beats all compared methods, serving as
benchmark for RAG evaluation in equity investment, integrating a strong baseline on FinTMMBench.
diverse data types for comprehensive analysis. As shown in Figure 1
(a), financial tables and news articles are simultaneously used for 2 Proposed FinTMMBench
calculating key financial ratios and assessing market sentiment in
fundamental analysis, and stock prices and technical charts are both Our FinTMMBench is constructed following a template-guided
required for calculating moving averages and identifying trends generation pipeline, as shown in Figure 2.
in technical analysis. Furthermore, equity analysis often involves
temporal-aware queries, which require precise identification of 2.1 Heterogeneous Corpus Preparation
time-specific information(e.g. year, month). For example, as shown To construct FinTMMBench, we collect financial data of the NASDAQ-
in Figure 1 (b), answering “What is the Price-to-Earnings (P/E) ratio 100 companies in 2022, which include four types as below.
of Apple on Dec 30, 2022, given 1,000,000 shares?” requires extracting
data for “Dec 30, 2022” from tables and stock prices, highlighting • Financial Tables:.For each company, we collect 12 quarterly
the need for temporal awareness. and 3 annual financial tables from 2022 via public APIs2 , totaling
To construct FinTMMBench, we collect 2022 financial data for 1,500 financial tables.
all NASDAQ-100 companies across four modalities. Working with fi- • News Articles: We gather over 70, 000 Reuters financial news ar-
nancial experts, we use a template-based approach to automatically ticles from 2021–2022, then filter for strong relevance to NASDAQ-
generate QA pairs, reflecting real-world analysis needs. About 100 100 companies, resulting in about 3, 100 articles.
templates with Chain-of-Thought (CoT) guideline cover 10 finan- • Daily Stock Prices: For each company, we collect 252 daily
cial tasks, such as information extraction, trend analysis and event records (high, low, open, close, volume) for 2022, totaling 25, 200
detection. Automatic revision and human review further enhance records. In total, we obtain 25, 200 records for all the companies.
data quality. In total, FinTMMBench contains 5, 676 high-quality • Visual Technical Charts: Weekly and monthly candlestick
QA pairs and 36, 100 raw data items. charts are generated from the daily stock price data.
Existing RAG methods, such as GraphRAG [8] and LightRAG [12], All data are standardized into JSON files, each storing a gran-
tend to struggle with answering the temporal-aware questions ular data point(e.g., a news article about a company), ensuring
across multi-modal financial data in our FinTMMBench, as shown consistency and seamless integration across modalities.
in Table 4. To address the challenge in FinTMMBench, we propose
a novel TMMHybridRAG method by combining dense retrieval
and graph retrieval techniques. First, TMMHybridRAG extracts
entities and their relations from each financial news article and 2 [Link]
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Step 2: Template and CoT Guidelines Design Step 3: QA Pairs Generation

CoT Multi-Modal LLM


Template Guidelines
Financial
QA Pairs
Experts
CoT
Template
Guidelines

Financial Financial Stock Technical Automatic Human


Tables News Prices Charts Revision Review

Step 1: Heterogeneous Corpus Preparation Step 4: Data Quality Assurance


Figure 2: An overall pipeline for constructing FinTMMBench.
SYSTEM_Prompt: You are a financial assistant... Table 1: Statistics of FinTMMBench.
Template: What is the P/B ratio of [Company] on [date], assuming that the
outstanding shares is [X]?
CoT Guideline:
1. Extract the totalShareholderEquity and price of [Company] on [date] Statistic Number
2. Book Value per Share = totalShareholderEquity / outstanding shares Total Number of Companies 100
3. P/B ratio = stock price / Book Value per Share
Data Points: Total Number of Raw Data 36,100
Close Price on Dec 29 2022: 129.61 (Source ID: 3e2f...)
Close Price on Dec 30 2022: 129.93 (Source ID: 59de...) # Financial Tables 1,500
Close Price on Jan 03 2023: 125.07 (Source ID: 64ef...)
Net Income in Dec 2022: 29,998.00 (Source ID: d9b2...)
# News Articles 3,133
... # Daily Stock Price 25,200
3-shot Example:
... INPUT # Visual Technical Charts 6,267
Multi-Modal LLM Total Number of Questions 5,676
Question: What is the Price-to-Earnings (P/E) ratio of Apple on Dec. 30, 2022, given Avg. Number of Question per Company 56.76
1,000,000 shares? Avg. Number of Words per Question 18.88
Explanation:
1. Extract the necessary information, close price is 129.93, Net Income is 29,998.00. Avg. Number of Words per Answer 6.87
2. Calculate the Earnings per Share, EPS= Net Income/shares= 0.029998
(29,998.00/ 1,000,000=0.029998)
3. Calculate the P/E ratio, P/E ratio=close price/EPS =4331.29
(129.93/0.029998= 4331.29)
Answer: 4,331.29 • Counterfactual Reasoning (CR): The question requires coun-
Source IDs: StockPrice-59de..., FinancialTable-d9b2..., ... OUTPUT terfactual reasoning to answer.
Figure 3: An example for QA pair generation. • Comparison (CP): The question requires comparing indicators
across different companies to obtain the answer.
• Sorting (ST): The question requires sorting indicators to infer
2.2 Template and CoT Guidelines Design
the answer.
Considering the high cost of human annotation, we design diverse • Counting (CT): The question requires counting the number of
question templates and corresponding CoT guidelines, to guide data points to infer the answer.
multi-modal LLMs to generate high-quality QA pairs automatically.
Specifically, we collaborate with financial experts to curate a set All questions are temporal-aware, requiring information from
of approximately 100 different question templates, which cover specific periods (e.g., day, month), and may encompass multiple
various financial tasks including: financial tasks. Detailed CoT guidelines for each template encourage
step-by-step reasoning, reducing inconsistencies in generated QA
• Information Extraction (IE): The question requires querying pairs and improving the quality of the dataset.
specific information (e.g., total revenue and net income) from the
financial corpus.
• Arithmetic Calculation (AC): The question requires deriving 2.3 QA Pair Generation
an indicator using a given formula based on relevant information. We employ GPT-4o-mini as the multi-modal LLM for QA pair gener-
• Trend Analysis (TA): The question requires analyzing the trend ation. As shown in Figure 3, the multi-modal LLM receives three key
of an indicator over time. inputs: 1) a question template, 2) a CoT guideline, and 3) some data
• Logical Reasoning (LR): The question requires logical reason- points of daily stock prices. Only data points relevant to each ques-
ing to infer the answer. tion template (e.g., news articles for event detection) are provided
• Sentiment Classification (SC): The question requires analyzing as input to the multi-modal LLM. This allows the multi-modal LLM
the sentiment polarity of a news article relevant to a specific to focus on essential information for QA generation. We prompt
company’s aspect (e.g., product and service). it to generate a question, step-by-step reasoning, the final answer,
• Event Detection (ED): The question requires identifying the and the IDs of referenced data points. Few-shot prompting is used
events mentioned in a news article. to further improve QA quality.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.

Table 2: Financial task distribution across different modali- Table 3: Comparison between our FinTMMBench with other
ties in FinTMMBench. Financial QA Datasets.

FA Task Table News Price Chart Hybrid Temporal Corpus Modality


Dataset RAG
Information Extraction 1,950 0 1,315 0 416 Question Tabular Textual Time-Series Visual
Arithmetic Calculation 1,494 0 1,112 0 416 FiQA-SA [18] ✗ ✗ ✗ ✓ ✗ ✗
Trend Analysis 575 0 489 421 0 FPB [19] ✗ ✗ ✗ ✓ ✗ ✗
Logical Reasoning 661 0 121 0 164 TAT-QA [32] ✗ ✗ ✓ ✓ ✗ ✗
Sentiment Classification 0 977 0 0 0 TAT-HQA [16] ✗ ✗ ✓ ✓ ✗ ✗
Event Detection 0 597 0 0 0 FinQA [4] ✗ ✗ ✓ ✓ ✗ ✗
Counterfactual Reasoning 778 0 539 0 416 MultiHiertt [29] ✗ ✗ ✓ ✓ ✗ ✗
Comparison 474 0 673 0 302 FinBen [28] ✗ ✗ ✗ ✓ ✗ ✗
Sorting 560 0 166 0 95 TAT-DQA [31] ✗ ✗ ✓ ✓ ✗ ✓
Counting 123 0 0 0 0 MultiModalQA [25] ✗ ✗ ✓ ✓ ✗ ✓
TempQuestions [14] ✗ ✓ ✗ ✓ ✓ ✗
AlphaFin [17] ✓ ✓ ✗ ✓ ✓ ✗
FinTextQA [3] ✓ ✗ ✗ ✓ ✗ ✗
2.4 Data Quality Assurance OmniEval [26] ✓ ✗ ✗ ✓ ✗ ✗
Automatic Revision. We preform automatic revision and human FinanceBench [13] ✓ ✗ ✗ ✓ ✗ ✓
FinTMMBench ✓ ✓ ✓ ✓ ✓ ✓
review to ensure the data quality of FinTMMBench. Specifically,
we develop a script to automatically check and revise the generated
QA pairs based on predefined rules. To name a few, the IDs of our FinTMMBench is designed to evaluate RAG systems in an-
the referred data points must be correct; the equations in each swering temporal-aware questions across a multi-modal corpus,
reasoning step must maintain equality between the left and right encompassing tabular, textual, time-series, and visual data.
sides; the answer inferred based on all reasoning steps must be
consistent with the final answer. 3 Proposed TMMHybridRAG Method
Human Review. After each round of automatic revision, we ran- To address the temporal-aware questions over heterogeneous fi-
domly select a set of samples based on the distribution of financial nancial data in FinTMMBench, we propose a novel RAG method
tasks and have two domain experts evaluate their accuracy, docu- TMMHybridRAG, which combines the dense and graph retrieval
menting any issues, which then inform the next round of automatic techniques, as shown in Figure 4.
revision. The verification results are then reviewed by a third expert
for additional validation. We repeat this iterative revision-review 3.1 Preprocessing
process until the verification accuracy exceeds 85% and the inter-
annotator agreement between the two experts reaches 85%. We generate textual descriptions for all non-textual data and then
identify entities and their relationships across different modalities
as preprocessing. In particular,
2.5 Dataset Analysis
• Financial Tables: Each financial table is treated as an entity with
As shown in Table 1, FinTMMBench consists of 34, 815 raw data
the temporal information determined by the period involved in
entries from NASDAQ-100 companies across four modalities, includ-
the table, and its name involves the company name, table name,
ing 1, 500 financial tables, 3, 133 news articles, 25, 200 daily stock
and the period described in this table. A summary of table is
price records, and 6, 267 visual technical charts. A total of 5, 676
generated by an LLM, serving as the entity’s description.
QA pairs are generated based on these raw data and CoT templates,
• News Articles: The enenties and relationships with their de-
with average length of the questions is 18.88 words, and the average
scriptions are directly extracted from each news article using an
length of the answers is 6.87 words.
LLM. The temporal information of an entity or relationship is
Table 2 shows the distribution of financial tasks across data
the publication date of the news article.
modalities. The questions in FinTMMBench span a wide range of
• Daily Stock Prices: Each daily stock price record is treated as a
financial tasks, enabling comprehensive evaluation of RAG systems
unique entity, named with the stock symbol and date. An LLM
on heterogeneous financial data.
generates a description for each record, and the date serves as its
temporal information. We link records from consecutive business
2.6 Comparison with Other Benchmarks days for the same company to capture the temporal information.
We further provide a comparison of our FinTMMBench with exist- • Visual Technical Charts: Each chart is regarded as an entity,
ing financial QA datasets to stress its merits, as shown in Table 3. with its name incorporating metadata like the company name,
It can be seen that most existing financial QA datasets are not and the time period represented in the chart. Then, we utilize a
open-domain, except for FinTextQA [3], FinanceBench [13], Al- multi-modal LLM to generate a concise summary for each chart,
phaFin [17] and OmniEval [26]. With the exceptions of TempQues- which serves as the entity’s description. The period depicted in
tions [14] and AlphaFin [17], few of them are designed to address the chart serves as the temporal information for the entity.
temporal-aware questions. In addition, existing datasets are mostly • Cross-Modality Relationships: Cross-modality relationships
restricted to specific modalities, such as textual data only (e.g., Fin- play a critical role in unifying the diverse data sources within
TextQA [3]), time-series data only (e.g., TempQuestions [14]), both the temporal knowledge graph. Specifically, we employ a multi-
tabular and textual data (e.g., TAT-QA [32]), or tabular, textual, modal LLM to automatically establish relationships across differ-
and visual data (e.g., MultiModalQA [25]). Compared with them, ent modalities by providing it the contextual information about
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Question: What is the Price-to-Earnings (P/E) ratio of Answer: 4,331.10


Apple on Dec. 30, 2022, given 1,000,000 shares? Explanation: The P/E ratio is calculated by dividing the stock price of
$129.93 by the earnings per share of $3.
Multi-Modal LLM

Query Keywords Retrieval Entity Multi-Modal LLM


Indexing Relation

Temporal-aware Dense Vector Retrieved Information


Company: Apple Company: Income Statement Table: (Company: Apple - Stock
Apple Apple Q4 2022 Price: Apple Dec 30 2022)

Stock Price: Apple Dec 30 2022

Temporal-aware Heterogeneous Graph


Income Statement Table: Apple

Raw Data Mapping


Q4 2022 Company:
Apple
(Company: Apple - Stock (Company: Apple - Income
(Company: Apple - Stock Price: Price: Apple Dec 30 2022) Statement Table: Apple Q4 2022)
Apple Dec 30 2022)
Stock Price: Income Statement Table:
(Company: Apple - Income
Apple Dec 30 2022 Apple Q4 2022
Statement Table: Apple Q4 2022)

Entities & Relations


Company: Stock Price: Income Statement (Company: Apple - Income (Company: Apple - Stock
Apple Apple Dec 30 2022 Table: Apple Q4 2022 Statement Table: Apple Q4 2022) Price: Apple Dec 30 2022)

Multi-Modal LLM

Heterogeneous Financial Corpus


Financial Financial Stock Technical
Tables News Prices Charts

Figure 4: Illustration of proposed TMMHybridRAG, a novel Temporal-Aware Multi-Modal RAG method.


the entities, including their names, associated metadata, and tex- 3.3 Retrieval
tual descriptions. With this information, the MLLM infers and We integrate dense and graph retrieval for enhanced effectiveness.
generates cross-modality relationships by identifying logical con- Keywords Identification and Expansion. Given a question, we
nections between the entities. first use an LLM to extract and expand relevant keywords, following
an approach similar to LightRAG [12]. These keywords, including
3.2 Indexing both entity and relationship names, are utilized to retrieve relevant
Temporal-aware Dense Vectors. First, TMMHybridRAG encodes entities and relationships from the dense vectors and graph.
each entity and relationship with its temporal information to gen- Dense Retrieval. We encode each query keyword into a dense
erate a dense vector using OpenAI embedding models (i.e., OpenAI vector and retrieve the top 𝐾 vectors from the vector database. Each
text-embedding-3-small). Then, we store all obtained dense vectors dense vector represents an entity or a relation.
in a vector database for further usage in the retrieval phase. Embed- Graph Retrieval. First, we aggregate all query keywords along
ding temporal information directly into the hidden representations with entity and relationship names obtained from dense retrieval.
allows for the retrieval of relevant entities and relationships based We then use these combined keywords to apply graph retrieval,
on their associated date period. searching for associated entities and relationships within the graph.
Temporal-aware Heterogeneous Graph. Knowledge graphs [11] Finally, all retrieved entities and relationships from both the vector
are powerful tools for representing relationships between diverse and the graph are utilized to generate the answer in the next step.
entities. TMMHybridRAG builds a knowledge graph with extracted
entities, e.g. company, person and location, and their relations with
an online LLM (i.e., GPT-4o-mini). Given the importance of temporal 3.4 Generation
information in the finance domain, each entity and relationship is With the retrieved entities and relations, we leverage a multi-modal
designed to store its corresponding temporal information as one of LLM to generate the final answer.
its properties. Additionally, each entity and relationship includes a Raw Data Mapping. First, we gather the raw data from different
textual description property and a Source ID attribute that facilitates modalities linked to the retrieved entities and relationships based
raw data mapping during the generation phase. on the source IDs. Although we generate a textual description for
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.

Table 4: Performance comparison between our TMMHy-


bridRAG and other baseline methods. Best and second-best No Retrieval setting, GPT-4o-mini performs poorly, revealing the
results are marked in bold and underlined, respectively. necessity of the RAG for correctly answering the questions in our
Setting Model EM (%) F1 Score Acc (%) LLM Acc (%) FinTMMBench. 2) Among all methods in No Visual setting, BGE-
No Retrieval GPT-4o-mini 4.51 5.89 6.68 6.25 Text achieves the highest scores compared to other methods. Our
BM25 10.85 20.89 15.89 11.97 TMMHybridRAG (No Visual) ranks the second and reaches compa-
Naive RAG 6.59 17.05 10.53 8.78 rable performance on all four metrics. 3) Our TMMHybridRAG (All)
GraphRAG 0.05 12.86 18.57 7.01
No Visual
LightRAG 4.62 15.07 8.32 8.32 consistently achieves the best results across all evaluation metrics,
BGE-Text 17.11 27.36 23.41 18.50 demonstrating the superiority of our method in addressing the
TMMHybridRAG 15.45 26.48 22.14 17.02
problems in FinTMMBench. Specifically, it attains an EM score of
CLIP-B 12.12 20.30 19.75 14.33
All 19.12%, an F1 score of 31.41, an accuracy of 26.56%, and an LLM-
BLIP-B 13.21 22.56 20.71 15.14
BGE-Visual 14.51 25.45 21.89 16.04 judge accuracy of 21.53%. 4) Though our TMMHybridRAG (All)
TMMHybridRAG 19.12 31.41 26.56 21.53
achieves state-of-the-art on FinTMMBench, the F1 score remains
each entity and relation, some crucial information or metrics may relatively low at 31.41. This highlights the significant challenges
be inadvertently lost without the raw data. By providing original inherent in FinTMMBench, demanding the development of more
data sources, we ensure that any analysis conducted is based on advanced RAG methods.
the correct and complete information.
Answer Generation. A multi-modal LLM is utilized to generate 4.3 In-Depth Analysis
the final answer, taking as input the question, the retrieved entities We further investigate the performance of methods across various
and relationships along with their temporal properties and textual financial tasks and data modalities. See results in Figure 5.
descriptions, and the corresponding raw data. The multi-modal Performance Analysis on Different Financial Tasks. As shown
LLM is instructed to output the intermediate reasoning steps and in Figure 5 (a), we can observe: 1) Our TMMHybridRAG (All) sig-
the final answer based on the multi-modal inputs. nificantly outperforms all other methods on most financial tasks,
demonstrating consistent effectiveness across diverse challenges
4 Experiments in the finance domain. 2) For Sentiment Classification which is
4.1 Experimental Settings designed to inquire about specific aspects of a company, requiring
Compared Methods. We employ three experimental settings. 1) the aggregation of dispersed information, GraphRAG achieves the
No Retrieval: No data retrieval is applied, and only the question best performance, possibly because its explicit high-level structures,
itself is fed into a multi-modal LLM to infer the answer; GPT-4o- like communities, can particularly benefit the summarization-based
mini is adopted in this setting. 2) No Visual: The retrieved tables, reasoning tasks. 3) For the Event Detection task, BGE-Visual ob-
news, stock prices and textual description of charts are fed into tains the highest performance, demonstrating the effectiveness of
a multi-modal LLM. BM25 [23], Naive RAG [10], GraphRAG [8], BGE-based models in processing textual news data.
LightRAG [12], and BGE-Text [27] are applied in this setting. 3) All: Performance Analysis Across Different Modalities. We present
All retrieved tables, news, stock prices, textual description of charts the performance of all methods across different modalities in Fig-
and the visual chart itself are used as the input of a multi-modal ure 5 (b). We find: 1) TMMHybridRAG (All) consistently beats all
LLM to derive the answer. BGE-Visual [30] is used in this setting. other methods across all modalities on our FinTMMBench, under-
Evaluation Metrics. Following the standard evaluation protocol, scoring its superiority in answering temporal-aware questions over
we use Exact Match (EM), F1 Score, and Accuracy (Acc) as evalua- multi-modal data. 2) Comparably, TMMHybridRAG is especially
tion metrics [22]. Additionally, to achieve a comprehensive assess- effective for questions involving visual technical charts, validating
ment of model performance, we employ LLMs as automated judges our approach to handling visual data through textual descriptions,
to assess model predictions compared to ground-truth answers. temporal information, and raw images. 3) In contrast, questions
Implementation Details. GPT-4o-mini is used to generate the that rely on multiple modalities and tabular data pose the great-
textual description in graph construction, and keywords in retrieval. est challenge for the TMMHybridRAG method, highlighting the
We use text-embedding-3-small to transform text chunks to dense difficulties of our FinTMMBench.
vectors. GPT-4o-mini is also used as the LLM evaluator. We use
Milvus as the vector database and neo4j as the graph database. GPT- 4.4 Ablation Study
4o-mini is applied as the multi-modal LLM to take the question and We conduct ablation study to evaluate effects of design choices in
the retrieved results as input to infer the answer. For BGE-Text and TMMHybridRAG, including temporal-aware dense vector, temporal-
BGE-Visual, we apply bge-large-en-v1.5 and bge-visualized-base- aware heterogeneous graph, raw data mapping, and incorporation
en-v1.5; for CLIP-B and BLIP-B, we use clip-vit-base-patch16 and of temporal information as properties in entities and relationships.
blip-image-captioning-base. See experiment results in Table 5.
• Removing Temporal-aware Dense Vectors (- Vec). In this
4.2 Main Results variant, the temporal-aware dense vector is removed. Given a
To verify the effectiveness of the proposed TMMHybridRAG, we query, the model searches for the entities and relationships from
compare its performance with baseline methods on the newly con- the graph only. This leads to a significant decline in performance
structed FinTMMBench. Experiment results are summarized in across all four evaluation metrics, e.g. 31.41 down to 11.45 for
Table 4, from which we make several key observations: 1) Under F1 score. The most substantial performance drop is observed on
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

Figure 5: Performance analysis on different financial tasks and modalities.


Table 5: Ablation study. Best and second-best results are marked in bold and underlined, respectively.
LLM-judge Acc (%) on Different Financial Tasks
Model EM (%) F1 Score Acc (%) LLM-judge Acc (%)
IE AC TA LR SC ED CR CP ST CT
TMMHybridRAG (All) 19.12 31.41 21.53 26.56 14.19 14.57 19.90 15.23 14.74 11.38 15.86 17.02 59.77 51.01
- Vec 6.34 11.45 6.32 7.26 9.74 5.79 13.89 11.31 12.35 4.88 8.70 12.89 1.54 5.90
- Graph 14.96 25.75 16.23 20.23 12.04 12.76 14.52 13.05 14.49 9.76 13.02 14.21 36.42 44.20
- Raw 15.63 27.58 16.75 26.44 20.38 27.40 26.03 20.55 14.98 12.20 15.47 17.75 38.41 40.44
- Temporal 17.16 28.38 19.29 23.70 12.84 10.04 18.83 14.24 13.64 13.93 12.81 16.16 56.28 50.25

the Sentiment Classification and Event Detection tasks over news 4.5 Performance Analysis on Different
articles. This reveals the importance of constructing dense vectors Multi-modal LLMs
for effectively addressing questions that depend on textual data.
We replace the multi-modal LLM used for answer generation with
• Removing Temporal-aware Heterogeneous Graph (- Graph).
other multi-modal LLMs and compare their performance. Com-
This variant removes the temporal-aware heterogeneous graph.
pared models are from different model families, including GPT-4o-
Given a query, all relevant entities and relationships are retrieved
mini [1], Llama 3.2 series [7], Qwen series [2], DeepSeek series[6],
from the temporal-aware dense vectors. A significant perfor-
and Gemini series [5], and Gemini series [5]. In Table 6 we summa-
mance drop across all four metrics can be observed. As Trend
rize parameter sizes, multi-modal LLMs, and their corresponding
Analysis requires understanding sequential relationships, the
performance on FinTMMBench. It can be seen that Gemini-2.0-
absence of the graph leads to worse performance. Note, the per-
Flash achieves the highest accuracy of 18.42%, followed by Kimi-VL-
formance on some tasks, including Sentiment Analysis, Logical
A3B-Instruct at 17.63%, surpassing both DeepSeek series and Qwen
Reasoning and Counting, is slightly better than the full model.
series. This suggests that even with the closed source models still
This may be because graph retrieval can introduce noise, hinder-
leading the pack, some open-source models can achieve competitive
ing the multi-modal LLM from identifying correct information.
performance. It also suggests that model size is not everything, in-
• Removing Raw Data Mapping (- Raw). This variant chooses
dicating that TMMHybridRAG, with its efficient architectures and
not to use raw data during answer generation, relying only on
techniques, does not rely on high-performance multi-modal LLMs
the retrieved entity and their relationships, which leads to a
to deliver competitive results. These results further demonstrate
noticeable drop across all metrics. For some tasks, e.g. Arithmetic
the broad applicability and effectiveness of our approach across
Calculation and Logical Reasoning, the performance is better than
diverse model classes and settings.
the full model. This may be because all necessary information for
answering the questions is already contained within the entities
4.6 Error Analysis
or relations, and raw data tends to include irrelevant details
misleading the multi-modal LLM in answer generation. We analyze error cases to better reveal the limitations of our TMMHy-
• Removing Temporal Information (- Temporal). This variant bridRAG and the challenges inherent in FinTMMBench. We ran-
removes temporal-related properties from all entities and rela- domly select 200 incorrect predictions and categorize the errors
tions, leading to worse performance than the full model across into four groups, as shown in Table 7, each with a representative ex-
all four metrics. The decline is especially obvious on Arithmetic ample. 1) Retrieval Error (46.5%): The retrieved data does not contain
Calculation and Trend Analysis tasks, highlighting the importance the key entities, relations, or relevant information needed to an-
of incorporating temporal information for effectively analyzing swer the question. 2) Calculation Error(29.0%): The model correctly
temporal-aware calculation and trend analysis in RAG systems. selects the relevant formula but makes mistakes in computation.
3) Reasoning Error (13.5%): The model misunderstands financial
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.

Table 6: Performance comparison of different multi-modal Table 7: Error Analysis. Q, G, P denote question, golden an-
LLMs with retrieval. swer, and TMMHybridRAG generated answer, respectively.
Q: What was CoStar Group’s otherCurrentAssets value
on March 31, 2022?
Model Open/Closed Params (B) Accuracy (%) Retrieval Error G: USD 36,183,000
(46.5%)
GPT-4o-mini Closed-source – 21.53 P: The retrieved tables do not contain any data the
Gemini-2.0-Flash Closed-source – 18.42 otherCurrentAssets value.
Kimi-VL-A3B-Instruct Open-source 16 17.63 Q: If Datadog had 15,000,000 shares instead of
Qwen2.5-7B-Instruct Open-source 7 15.15 10,000,000 and a book value of USD 2,000,000,000 ,
DeepSeek R1 8B Open-source 8 13.41 what would its P/B ratio be on Jan 5, 2022?
Calculation Error
DeepSeek R1 14B Open-source 14 13.15 (29.0%) G: 1.036
Llama 3.2 11B Open-source 11 11.57 P: Book Value per Share: 2,000,000,000
10,000,000 = 20
Llama 3.2 3B Open-source 3 11.27 Q: If Ansys’s stock price trend from October 13, 2022,
Qwen-VL-Chat Open-source 7 8.89 continued, what would its price be next month?
Reasoning Error G: 207.68 * (1 + 0.0769) = USD 223.66
concepts, misinterprets relationships between variables, or applies (13.5%) P: With the price reaching a last closing price of USD
incorrect logical reasoning to infer the answer. 4) Temporal Error 279.21 ...
(5.5%): The model uses data from the correct source but associates Q: When did AEP experience the lowest price in Sep-
it with the wrong timestamp. Temporal Error tember 2022?
We make following observations: 1) Most errors are Retrieval Er- (5.5%) G: September 30, 2022
rors (46.5%). This suggests advanced indexing or retrieval methods P: On October 29, 2022, the stock ...

are demanded to improve recall in information retrieval. 2) Chal-


temporal-aware multi-modal benchmark designed to evaluate RAG
lenges persist in Arithmetic Computation and Complex Reasoning.
systems in finance. It encompasses financial data across four modal-
Calculation Errors and Reasoning Errors collectively account for
ities—tabular, textual, time-series, and visual data. Additionally,
42.5% of failures, underscoring the challenges multi-modal LLMs
all questions in FinTMMBench are temporal-aware, addressing a
face in performing arithmetic computations and complex reasoning.
critical gap in existing benchmarks.
To address these issues, two possible approaches can be consid-
ered: i) to improve quality of retained information in the retrieval, 5.2 Graph-based RAG
such as reducing irrelevant content; ii) to enhance LLMs’ under-
standing of financial terminology, improve their ability to perform RAG [15, 33] has been widely used to enhance performance of
complex financial reasoning, and integrate external tools to assist LLMs across various tasks by integrating an Information Retriever
with numerical computations. iii) Temporal Inference is crucial. (IR) module to leverage external knowledge. Recently, graph-based
Though less frequent, Temporal Errors (5.5%) are unignorable for RAG methods [8, 12, 24, 34, 35] have demonstrated remarkable per-
time-sensitive tasks, as incorrect temporal inference can result in formance across diverse applications. For instance, GraphRAG [8]
significant factual inaccuracies. improves traditional RAG by building a knowledge graph from
extracted entities and relations, grouping related entities into com-
munities, and generating summaries for each. During inference, it
5 Related Work synthesizes answers from these community summaries. Hybrid [24]
5.1 Financial QA Datasets and LightRAG [12] enhance GraphRAG by combining dense re-
To date, many financial QA datasets have been released to advance trieval with graph retrieval techniques. Despite effectiveness, all
research in financial analysis, which can be divided to Non-RAG these methods primarily focus on textual data, resulting in sub-
QA, Text-RAG QA, and Multi-Modal-RAG QA datasets. Non-RAG optimal performance when handling multi-modal data. Moreover,
QA [18, 28, 32] datasets focus on financial analysis using relatively they struggle to effectively address temporal-aware queries in FinT-
short context information that can be directly input into LLMs. For MMBench. We propose TMMHybridRAG, a novel graph-based
example, FiQA-SA [18] and FPB [19] are designed for emotion anal- RAG approach specifically designed to tackle the challenges of
ysis based on financial texts; TAT-QA [32] and FinQA [4] aim to an- temporal-aware multi-modal RAG presented in FinTMMBench.
swer questions given a financial table and its associated paragraphs
extracted from financial reports. Text-RAG QA datasets, e.g. Fin- 6 Conclusion
TextQA [3] and OmniEval [26], are aimed at evaluating text-based In this work, we introduce FinTMMBench, the first benchmark for
RAG systems in finance. For instance, FinTextQA [3] is a long-form evaluating temporal-aware multi-modal Retrieval-Augmented Gen-
QA dataset containing 1,262 high-quality QA pairs that require eration (RAG) systems in financial analysis. FinTMMBench com-
RAG systems to address based on finance textbooks and policy and prises 5,676 questions spanning financial tables, news articles, stock
regulation from government agency websites. Current Multi-Modal prices, and technical charts, designed to assess a model’s ability to
RAG QA datasets include FinanceBench [13], incorporating time- retrieve and reason over temporal financial information. To address
series data in addition to textual data, and AlphaFin [17], involving its challenges, we propose TMMHybridRAG, a novel approach
visual data with textual data to assess RAG systems. Though with integrating dense and graph retrieval with temporal-aware entity
notable strengths, these datasets are limited to specific modalities, modeling. Our experiments show TMMHybridRAG outperforms
and only AlphaFin incorporates some temporal questions focused existing methods, yet the generally low performance also highlights
on time-series data. In comparison, our FinTMMBench is the first the persisting challenges of our FinTMMBench.
Towards Temporal-Aware Multi-Modal Retrieval Augmented Generation in Finance Conference acronym ’XX, June 03–05, 2018, Woodstock, NY

References [15] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin,
[1] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel,
cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for
Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing
(2023). Systems, Vol. 33. Curran Associates, Inc., 9459–9474. [Link]
[2] Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang cc/paper_files/paper/2020/file/[Link]
Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Frontier Large Vision- [16] Moxin Li, Fuli Feng, Hanwang Zhang, Xiangnan He, Fengbin Zhu, and Tat-Seng
Language Model with Versatile Abilities. arXiv preprint arXiv:2308.12966 (2023). Chua. 2022. Learning to imagine: Integrating counterfactual thinking in neural
[3] Jian Chen, Peilin Zhou, Yining Hua, Yingxin Loh, Kehui Chen, Ziyuan Li, Bing discrete reasoning. In Proceedings of the 60th Annual Meeting of the Association
Zhu, and Junwei Liang. 2024. FinTextQA: A Dataset for Long-form Financial for Computational Linguistics (Volume 1: Long Papers). 57–69.
Question Answering. arXiv preprint arXiv:2405.09980 (2024). [17] Xiang Li, Zhenyu Li, Chen Shi, Yong Xu, Qing Du, Mingkui Tan, Jun Huang,
[4] Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan and Wei Lin. 2024. AlphaFin: Benchmarking Financial Analysis with Retrieval-
Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, et al. Augmented Stock-Chain Framework. arXiv:2403.12582 [[Link]]
2021. Finqa: A dataset of numerical reasoning over financial data. arXiv preprint [18] Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott,
arXiv:2109.00122 (2021). Manel Zarrouk, and Alexandra Balahur. 2018. Www’18 open challenge: financial
[5] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen opinion mining and question answering. In Companion proceedings of the the web
Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, conference 2018. 1941–1942.
et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- [19] Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala.
modality, long context, and next generation agentic capabilities. arXiv preprint 2014. Good debt or bad debt: Detecting semantic orientations in economic texts.
arXiv:2507.06261 (2025). Journal of the Association for Information Science and Technology 65, 4 (2014),
[6] DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu 782–796.
Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang [20] Harry M Markowitz. 1991. Foundations of portfolio theory. The journal of finance
Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, 46, 2 (1991), 469–477.
Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, [21] Michael Power. 2004. The risk management of everything. The Journal of Risk
Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Deli Finance 5, 3 (2004), 58–65.
Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, [22] P Rajpurkar. 2016. Squad: 100,000+ questions for machine comprehension of text.
Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, arXiv preprint arXiv:1606.05250 (2016).
Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, [23] Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance
Jiawei Wang, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi framework: BM25 and beyond. Foundations and Trends® in Information Retrieval
Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, 3, 4 (2009), 333–389.
Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, [24] Bhaskarjit Sarmah, Dhagash Mehta, Benika Hall, Rohan Rao, Sunil Patel, and
Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Meng Li, Stefano Pasquali. 2024. HybridRAG: Integrating Knowledge Graphs and Vector
Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Retrieval Augmented Generation for Efficient Information Extraction. In Pro-
Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji ceedings of the 5th ACM International Conference on AI in Finance (Brooklyn, NY,
Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang USA) (ICAIF ’24). Association for Computing Machinery, New York, NY, USA,
Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. 608–616. doi:10.1145/3677052.3698671
Li, Shuang Zhou, Shaoqing Wu, Shengfeng Ye, Tao Yun, Tian Pei, Tianyu Sun, [25] Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai,
T. Wang, Wangding Zeng, Wanjia Zhao, Wen Liu, Wenfeng Liang, Wenjun Gao, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. [n. d.]. Multi-
Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, ModalQA: complex question answering over text, tables and images. In Interna-
Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu tional Conference on Learning Representations.
Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, [26] Shuting Wang, Jiejun Tan, Zhicheng Dou, and Ji-Rong Wen. 2024. OmniEval:
Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial
Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Domain. arXiv preprint arXiv:2412.13018 (2024).
Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, [27] Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023.
Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan C-Pack: Packaged Resources To Advance General Chinese Embedding.
Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, arXiv:2309.07597 [[Link]]
Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, [28] Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru
Yanhong Xu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, et al. 2024. The fin-
Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, ben: An holistic financial benchmark for large language models. arXiv preprint
Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang arXiv:2402.12659 (2024).
Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, [29] Yilun Zhao, Yunxiang Li, Chenying Li, and Rui Zhang. 2022. MultiHiertt: Numeri-
Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. 2025. cal Reasoning over Multi Hierarchical Tabular and Textual Data. In Proceedings of
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement the 60th Annual Meeting of the Association for Computational Linguistics (Volume
Learning. arXiv:2501.12948 [[Link]] [Link] 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.).
[7] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Association for Computational Linguistics, 6588–6600.
Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, [30] Junjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao, and Yongping Xiong. 2024. VISTA:
et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024). Visualized Text Embedding For Universal Multi-Modal Retrieval. arXiv preprint
[8] Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva arXiv:2406.04292 (2024).
Mody, Steven Truitt, and Jonathan Larson. 2024. From local to global: A graph [31] Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng
rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130 Chua. 2022. Towards complex document understanding by discrete reasoning. In
(2024). Proceedings of the 30th ACM International Conference on Multimedia. 4857–4866.
[9] Eugene F Fama and Kenneth R French. 1992. The cross-section of expected stock [32] Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang,
returns. the Journal of Finance 47, 2 (1992), 427–465. Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A Question An-
[10] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, swering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In
Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large Proceedings of the 59th Annual Meeting of the Association for Computational Lin-
language models: A survey. arXiv preprint arXiv:2312.10997 (2023). guistics and the 11th International Joint Conference on Natural Language Processing
[11] Google. 2012. Introducing the Knowledge Graph: Things, Not Strings. [Link] (Volume 1: Long Papers). 3277–3287.
google/products/search/introducing-knowledge-graph-things-not/ Accessed: [33] Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and
2025-01-07. Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-
[12] Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2024. Lightrag: domain question answering. arXiv preprint arXiv:2101.00774 (2021).
Simple and fast retrieval-augmented generation. (2024). [34] Fengbin Zhu, Moxin Li, Junbin Xiao, Fuli Feng, Chao Wang, and Tat Seng Chua.
[13] Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, 2023. Soargraph: Numerical reasoning over financial table-text data via semantic-
and Bertie Vidgen. 2023. Financebench: A new benchmark for financial question oriented hierarchical graphs. In Companion Proceedings of the ACM Web Confer-
answering. arXiv preprint arXiv:2311.11944 (2023). ence 2023. 1236–1244.
[14] Zhen Jia, Abdalghani Abujabal, Rishiraj Saha Roy, Jannik Strötgen, and Gerhard [35] Fengbin Zhu, Chao Wang, Fuli Feng, Zifeng Ren, Moxin Li, and Tat-Seng Chua.
Weikum. 2018. Tempquestions: A benchmark for temporal question answering. 2023. Doc2SoarGraph: Discrete reasoning over visually-rich table-text documents
In Companion Proceedings of the The Web Conference 2018. 1057–1062. via semantic-oriented hierarchical graphs. arXiv preprint arXiv:2305.01938 (2023).
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.

Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009

You might also like