Projects
Asset allocation from an LLM’s reflections on its own decisions
Market-Adaptive Memory LLM for the Black–Litterman Model
APIEMS 2025 oral, KIIE 2025 Spring Joint Conference poster · Code
Background
Mean-variance optimization is sensitive to errors in expected-return estimates, so the weights it produces tend to concentrate. The Black–Litterman model softens this by blending an investor’s views with market equilibrium returns, but forming those views and updating them in time is still a human job. I looked at whether that step can be automated using unstructured information such as news together with the record of how past decisions turned out.
Method
Every two weeks the LLM forms views and the Black–Litterman model blends them with equilibrium returns to set the weights.
- The LLM (GPT-4o-mini) reads recent prices, technical indicators, and sector news, and proposes an expected return for each stock. Sampling the same prompt several times gives the view uncertainty matrix Ω.
- When a rebalancing period underperforms, the agent looks back at the weights it held, the result, and the news at the time, and writes down what to change.
- These reflections go into a vector database. At the next decision the agent retrieves either the most recent ones or the ones from market regimes similar to today.
- The LLM views, which may be biased, are combined with equilibrium returns inside the Black–Litterman posterior, and mean-variance optimization turns that into the final weights.
An actual reflection the agent wrote:
The worst rebalance over-allocated to tech names (ADBE, GOOGL, META, V) even as sector indicators weakened. Overconfidence and a high risk-aversion setting amplified the drawdown (−2.63%). Before each trade, confirm momentum with at least two indicators such as RSI and MACD, and cut the risk-aversion coefficient by 20% when realized volatility exceeds 15%.
Experiment
The backtest runs from September 2024 to March 2025, a window chosen to fall after GPT-4o-mini’s knowledge cutoff. The universe is the five largest stocks by market capitalization in each of the 11 GICS sectors (55 in total), rebalanced every two weeks, long-only, with a 20% cap per name. I compared four settings: no reflection, recent reflections only, similar-regime reflections only, and both.
Results
- Retrieving only the similar-regime reflections worked best, ahead of the equal-weight and cap-weighted portfolios built from the same 55 stocks. What mattered was not the freshest memory but the one from a market that resembled the present.
- Using both the recent and the similar-regime reflections was the worst setting and trailed the benchmarks. With too much context, the signals that mattered appear to get diluted.
Preserving low-variance pricing information in vector-quantized factor models (in progress)
Applying vector quantization to an autoencoder-based factor model compresses the return structure into a small set of representative vectors. The problem is that even when the reconstruction error is small, directions that carry little variance but still matter for pricing can disappear.
I implemented a version that splits the latent space into subspaces and gives each one its own codebook, and I am checking on synthetic data how well low-variance factors survive. At an equal bit budget it improves the equal-weight portfolio over a single codebook, but the gain shrinks considerably under cap weighting. That suggests the improvement may lean on small caps, so I am running experiments that control for size.
Other work
NH Investment & Securities Big Data Competition, ETF curation (2024, Encouragement Award)
I used LangGraph and RAG to generate ETF reports grounded in news and company filings for each holding. So that the generated text was not taken at face value, I built a separate quantitative engine that scores ETFs on return, maximum drawdown, dividends, liquidity, and investor sentiment. Rather than relying on the return field supplied with the competition data, I crawled five years of adjusted closes from Yahoo Finance and recomputed one-month, three-month, and one-year returns, volatility, and drawdown myself, so that every ETF was measured on the same basis.
WorldQuant IQC 2024
I designed and backtested alpha ideas drawn from the literature on US and Chinese equity data, finishing first on campus and advancing to Stage 2.
UFEA financial engineering society (2024, vice president)
We worked through Wilmott’s Quantitative Finance and Derman’s The Volatility Smile, covering option pricing, volatility modeling, and hedging, and I implemented implied volatility calculations and the P&L of a delta-hedged portfolio. We then moved into fixed income and interest rate theory using Veronesi, Hull, and Neftci, studying how credit spreads and default risk are reflected in bond pricing, term structure models, and interest rate derivatives such as swaps, as well as the tranche structure and pricing logic behind structured credit products like CDOs. To connect this theoretical work to practical implementation, I also participated in a C/C++ study group within the club, working through John Armstrong’s C++ for Financial Mathematics and Mark Joshi’s C++ Design Patterns and Derivatives Pricing to implement option pricing logic and a multi-threaded Monte Carlo simulation engine in C++.