vix.ing · top · new · best · stats

Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking

2025/05/16 by C. Wei Li, Li, Changlun, Shi, Yao +14 · 5 citations
Computer Science · Decision Sciences · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Benchmarking #Computational Engineering #Earnings #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Finance #Financial Markets and Investment Strategies #Investment fund #Investment management #Investment strategy #Manager of managers fund #Market timing #Multiagent Systems (cs.MA) #Portfolio #Stock Market Forecasting Methods #Target date fund #and Science (cs.CE)

paper · pdf · doi:10.48550/arxiv.2505.11065

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/05/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Large Language Models (LLMs) have demonstrated notable capabilities across financial tasks, including financial report summarization, earnings call transcript analysis, and asset classification. However, their real-world effectiveness in managing complex fund investment remains inadequately assessed. A fundamental limitation of existing benchmarks for evaluating LLM-driven trading strategies is their reliance on historical back-testing, inadvertently enabling LLMs to "time travel"-leveraging future information embedded in their training corpora, thus resulting in possible information leakage and overly optimistic performance estimates. To address this issue, we introduce DeepFund, a live fund benchmark tool designed to rigorously evaluate LLM in real-time market conditions. Utilizing a multi-agent architecture, DeepFund connects directly with real-time stock market data-specifically data published after each model pretraining cutoff-to ensure fair and leakage-free evaluations. Empirical tests on nine flagship LLMs from leading global institutions across multiple investment dimensions-including ticker-level analysis, investment decision-making, portfolio management, and risk control-reveal significant practical challenges. Notably, even cutting-edge models such as DeepSeek-V3 and Claude-3.7-Sonnet incur net trading losses within DeepFund real-time evaluation environment, underscoring the present limitations of LLMs for active fund management. Our code is available at https://github.com/HKUSTDial/DeepFund.

Citations

Cited by

Related