Multimodal Rag
Multimodal Rag
Generation in Finance
Fengbin Zhu∗ Junfeng Li∗ Liangming Pan†
National University of Singapore National University of Singapore Peking University
Singapore Singapore China
zhfengbin@[Link] lijunfeng@[Link] peterpan10211020@[Link]
Buy, Sell or Hold? employs an LLM to generate descriptions for each entity and re-
lation. For non-textual data, TMMHybridRAG regards each table,
Fundamental Analysis Technical Analysis daily stock price record, and chart as a distinct entity and utilizes
• Debt-to-Equity Ratio • Relative Strength Index
• P/E Ratio • Moving Average an advanced multi-modal LLM to generate a textual summary for
• Market Sentiment • Trend
• ... • ... each, which serves as the entity’s description. Further, TMMHy-
bridRAG integrates temporal information into every entity and
Financial Financial Stock Technical relation as the properties to construct dense vectors and graphs.
Tables News Prices Charts
During prediction, given a question, all retrieved entities and re-
(a) Data-driven Equity Investment
lations from both dense vectors and graphs, along with their raw
Apple Inc. Income Statement Table Question: data, are fed into a multi-modal LLM to infer the answer. Extensive
Fiscal Period Dec 2022 Apr 2023 Jul 2023 What is the Price-to-Earnings
(P/E) ratio of Apple on Dec. experiments show that our TMMHybridRAG method significantly
Total Revenue 117,154.00 94,836.00 81,797.00
30, 2022, given 1,000,000 outperforms all compared methods across all evaluation metrics.
Gross Profit 50,332.00 41,976.00 36,413.00
shares?
Net Income 29,998.00 24,160.00 19,881.00 However, its F1 score remains relatively low at 31.41, highlight-
Operating Income 36,016.00 28,318.00 22,998.00 Supporting Evidence:
The close price of Apple on ing the substantial challenges presented in FinTMMBench and
Apple Inc. Stock Price
132
Dec. 30, 2022, is 129.93 USD. underscoring the need for more advanced RAG methods.
130 Date Close Price The net income of Apple in
128 2022Q4 is 29,998.00 USD. In summary, our major contributions are threefold: 1) To the best
126 Dec 29 2022 129.61
124
Dec 30 2022 129.93 Answer: 4331.29 of our knowledge, we are the first to investigate temporal-aware
122
Jan 03 2023 125.07 Explanation: P/E Ratio = multi-modal RAG in the financial domain, addressing a critical real-
23
23
23
2
2
02
02
02
02
20
20
20
82
92
02
5
c2
c2
c2
c3
n0
n0
n0
De
De
De
Ja
Ja
Ja
Table 2: Financial task distribution across different modali- Table 3: Comparison between our FinTMMBench with other
ties in FinTMMBench. Financial QA Datasets.
Multi-Modal LLM
the Sentiment Classification and Event Detection tasks over news 4.5 Performance Analysis on Different
articles. This reveals the importance of constructing dense vectors Multi-modal LLMs
for effectively addressing questions that depend on textual data.
We replace the multi-modal LLM used for answer generation with
• Removing Temporal-aware Heterogeneous Graph (- Graph).
other multi-modal LLMs and compare their performance. Com-
This variant removes the temporal-aware heterogeneous graph.
pared models are from different model families, including GPT-4o-
Given a query, all relevant entities and relationships are retrieved
mini [1], Llama 3.2 series [7], Qwen series [2], DeepSeek series[6],
from the temporal-aware dense vectors. A significant perfor-
and Gemini series [5], and Gemini series [5]. In Table 6 we summa-
mance drop across all four metrics can be observed. As Trend
rize parameter sizes, multi-modal LLMs, and their corresponding
Analysis requires understanding sequential relationships, the
performance on FinTMMBench. It can be seen that Gemini-2.0-
absence of the graph leads to worse performance. Note, the per-
Flash achieves the highest accuracy of 18.42%, followed by Kimi-VL-
formance on some tasks, including Sentiment Analysis, Logical
A3B-Instruct at 17.63%, surpassing both DeepSeek series and Qwen
Reasoning and Counting, is slightly better than the full model.
series. This suggests that even with the closed source models still
This may be because graph retrieval can introduce noise, hinder-
leading the pack, some open-source models can achieve competitive
ing the multi-modal LLM from identifying correct information.
performance. It also suggests that model size is not everything, in-
• Removing Raw Data Mapping (- Raw). This variant chooses
dicating that TMMHybridRAG, with its efficient architectures and
not to use raw data during answer generation, relying only on
techniques, does not rely on high-performance multi-modal LLMs
the retrieved entity and their relationships, which leads to a
to deliver competitive results. These results further demonstrate
noticeable drop across all metrics. For some tasks, e.g. Arithmetic
the broad applicability and effectiveness of our approach across
Calculation and Logical Reasoning, the performance is better than
diverse model classes and settings.
the full model. This may be because all necessary information for
answering the questions is already contained within the entities
4.6 Error Analysis
or relations, and raw data tends to include irrelevant details
misleading the multi-modal LLM in answer generation. We analyze error cases to better reveal the limitations of our TMMHy-
• Removing Temporal Information (- Temporal). This variant bridRAG and the challenges inherent in FinTMMBench. We ran-
removes temporal-related properties from all entities and rela- domly select 200 incorrect predictions and categorize the errors
tions, leading to worse performance than the full model across into four groups, as shown in Table 7, each with a representative ex-
all four metrics. The decline is especially obvious on Arithmetic ample. 1) Retrieval Error (46.5%): The retrieved data does not contain
Calculation and Trend Analysis tasks, highlighting the importance the key entities, relations, or relevant information needed to an-
of incorporating temporal information for effectively analyzing swer the question. 2) Calculation Error(29.0%): The model correctly
temporal-aware calculation and trend analysis in RAG systems. selects the relevant formula but makes mistakes in computation.
3) Reasoning Error (13.5%): The model misunderstands financial
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.
Table 6: Performance comparison of different multi-modal Table 7: Error Analysis. Q, G, P denote question, golden an-
LLMs with retrieval. swer, and TMMHybridRAG generated answer, respectively.
Q: What was CoStar Group’s otherCurrentAssets value
on March 31, 2022?
Model Open/Closed Params (B) Accuracy (%) Retrieval Error G: USD 36,183,000
(46.5%)
GPT-4o-mini Closed-source – 21.53 P: The retrieved tables do not contain any data the
Gemini-2.0-Flash Closed-source – 18.42 otherCurrentAssets value.
Kimi-VL-A3B-Instruct Open-source 16 17.63 Q: If Datadog had 15,000,000 shares instead of
Qwen2.5-7B-Instruct Open-source 7 15.15 10,000,000 and a book value of USD 2,000,000,000 ,
DeepSeek R1 8B Open-source 8 13.41 what would its P/B ratio be on Jan 5, 2022?
Calculation Error
DeepSeek R1 14B Open-source 14 13.15 (29.0%) G: 1.036
Llama 3.2 11B Open-source 11 11.57 P: Book Value per Share: 2,000,000,000
10,000,000 = 20
Llama 3.2 3B Open-source 3 11.27 Q: If Ansys’s stock price trend from October 13, 2022,
Qwen-VL-Chat Open-source 7 8.89 continued, what would its price be next month?
Reasoning Error G: 207.68 * (1 + 0.0769) = USD 223.66
concepts, misinterprets relationships between variables, or applies (13.5%) P: With the price reaching a last closing price of USD
incorrect logical reasoning to infer the answer. 4) Temporal Error 279.21 ...
(5.5%): The model uses data from the correct source but associates Q: When did AEP experience the lowest price in Sep-
it with the wrong timestamp. Temporal Error tember 2022?
We make following observations: 1) Most errors are Retrieval Er- (5.5%) G: September 30, 2022
rors (46.5%). This suggests advanced indexing or retrieval methods P: On October 29, 2022, the stock ...
References [15] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin,
[1] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel,
cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for
Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing
(2023). Systems, Vol. 33. Curran Associates, Inc., 9459–9474. [Link]
[2] Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang cc/paper_files/paper/2020/file/[Link]
Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Frontier Large Vision- [16] Moxin Li, Fuli Feng, Hanwang Zhang, Xiangnan He, Fengbin Zhu, and Tat-Seng
Language Model with Versatile Abilities. arXiv preprint arXiv:2308.12966 (2023). Chua. 2022. Learning to imagine: Integrating counterfactual thinking in neural
[3] Jian Chen, Peilin Zhou, Yining Hua, Yingxin Loh, Kehui Chen, Ziyuan Li, Bing discrete reasoning. In Proceedings of the 60th Annual Meeting of the Association
Zhu, and Junwei Liang. 2024. FinTextQA: A Dataset for Long-form Financial for Computational Linguistics (Volume 1: Long Papers). 57–69.
Question Answering. arXiv preprint arXiv:2405.09980 (2024). [17] Xiang Li, Zhenyu Li, Chen Shi, Yong Xu, Qing Du, Mingkui Tan, Jun Huang,
[4] Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan and Wei Lin. 2024. AlphaFin: Benchmarking Financial Analysis with Retrieval-
Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, et al. Augmented Stock-Chain Framework. arXiv:2403.12582 [[Link]]
2021. Finqa: A dataset of numerical reasoning over financial data. arXiv preprint [18] Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott,
arXiv:2109.00122 (2021). Manel Zarrouk, and Alexandra Balahur. 2018. Www’18 open challenge: financial
[5] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen opinion mining and question answering. In Companion proceedings of the the web
Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, conference 2018. 1941–1942.
et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multi- [19] Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala.
modality, long context, and next generation agentic capabilities. arXiv preprint 2014. Good debt or bad debt: Detecting semantic orientations in economic texts.
arXiv:2507.06261 (2025). Journal of the Association for Information Science and Technology 65, 4 (2014),
[6] DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu 782–796.
Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang [20] Harry M Markowitz. 1991. Foundations of portfolio theory. The journal of finance
Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, 46, 2 (1991), 469–477.
Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, [21] Michael Power. 2004. The risk management of everything. The Journal of Risk
Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Deli Finance 5, 3 (2004), 58–65.
Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, [22] P Rajpurkar. 2016. Squad: 100,000+ questions for machine comprehension of text.
Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, arXiv preprint arXiv:1606.05250 (2016).
Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, [23] Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance
Jiawei Wang, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi framework: BM25 and beyond. Foundations and Trends® in Information Retrieval
Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, 3, 4 (2009), 333–389.
Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, [24] Bhaskarjit Sarmah, Dhagash Mehta, Benika Hall, Rohan Rao, Sunil Patel, and
Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Meng Li, Stefano Pasquali. 2024. HybridRAG: Integrating Knowledge Graphs and Vector
Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Retrieval Augmented Generation for Efficient Information Extraction. In Pro-
Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji ceedings of the 5th ACM International Conference on AI in Finance (Brooklyn, NY,
Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang USA) (ICAIF ’24). Association for Computing Machinery, New York, NY, USA,
Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. 608–616. doi:10.1145/3677052.3698671
Li, Shuang Zhou, Shaoqing Wu, Shengfeng Ye, Tao Yun, Tian Pei, Tianyu Sun, [25] Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai,
T. Wang, Wangding Zeng, Wanjia Zhao, Wen Liu, Wenfeng Liang, Wenjun Gao, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. [n. d.]. Multi-
Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, ModalQA: complex question answering over text, tables and images. In Interna-
Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu tional Conference on Learning Representations.
Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, [26] Shuting Wang, Jiejun Tan, Zhicheng Dou, and Ji-Rong Wen. 2024. OmniEval:
Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial
Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Domain. arXiv preprint arXiv:2412.13018 (2024).
Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, [27] Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023.
Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan C-Pack: Packaged Resources To Advance General Chinese Embedding.
Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, arXiv:2309.07597 [[Link]]
Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, [28] Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru
Yanhong Xu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, et al. 2024. The fin-
Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, ben: An holistic financial benchmark for large language models. arXiv preprint
Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang arXiv:2402.12659 (2024).
Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, [29] Yilun Zhao, Yunxiang Li, Chenying Li, and Rui Zhang. 2022. MultiHiertt: Numeri-
Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. 2025. cal Reasoning over Multi Hierarchical Tabular and Textual Data. In Proceedings of
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement the 60th Annual Meeting of the Association for Computational Linguistics (Volume
Learning. arXiv:2501.12948 [[Link]] [Link] 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.).
[7] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Association for Computational Linguistics, 6588–6600.
Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, [30] Junjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao, and Yongping Xiong. 2024. VISTA:
et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024). Visualized Text Embedding For Universal Multi-Modal Retrieval. arXiv preprint
[8] Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva arXiv:2406.04292 (2024).
Mody, Steven Truitt, and Jonathan Larson. 2024. From local to global: A graph [31] Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng
rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130 Chua. 2022. Towards complex document understanding by discrete reasoning. In
(2024). Proceedings of the 30th ACM International Conference on Multimedia. 4857–4866.
[9] Eugene F Fama and Kenneth R French. 1992. The cross-section of expected stock [32] Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang,
returns. the Journal of Finance 47, 2 (1992), 427–465. Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A Question An-
[10] Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, swering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In
Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large Proceedings of the 59th Annual Meeting of the Association for Computational Lin-
language models: A survey. arXiv preprint arXiv:2312.10997 (2023). guistics and the 11th International Joint Conference on Natural Language Processing
[11] Google. 2012. Introducing the Knowledge Graph: Things, Not Strings. [Link] (Volume 1: Long Papers). 3277–3287.
google/products/search/introducing-knowledge-graph-things-not/ Accessed: [33] Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and
2025-01-07. Tat-Seng Chua. 2021. Retrieving and reading: A comprehensive survey on open-
[12] Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2024. Lightrag: domain question answering. arXiv preprint arXiv:2101.00774 (2021).
Simple and fast retrieval-augmented generation. (2024). [34] Fengbin Zhu, Moxin Li, Junbin Xiao, Fuli Feng, Chao Wang, and Tat Seng Chua.
[13] Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, 2023. Soargraph: Numerical reasoning over financial table-text data via semantic-
and Bertie Vidgen. 2023. Financebench: A new benchmark for financial question oriented hierarchical graphs. In Companion Proceedings of the ACM Web Confer-
answering. arXiv preprint arXiv:2311.11944 (2023). ence 2023. 1236–1244.
[14] Zhen Jia, Abdalghani Abujabal, Rishiraj Saha Roy, Jannik Strötgen, and Gerhard [35] Fengbin Zhu, Chao Wang, Fuli Feng, Zifeng Ren, Moxin Li, and Tat-Seng Chua.
Weikum. 2018. Tempquestions: A benchmark for temporal question answering. 2023. Doc2SoarGraph: Discrete reasoning over visually-rich table-text documents
In Companion Proceedings of the The Web Conference 2018. 1057–1062. via semantic-oriented hierarchical graphs. arXiv preprint arXiv:2305.01938 (2023).
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Tor et al.