Time-series foundation models for finance
A study of Amazon's Chronos-2, a pretrained time-series transformer, on Magnificent-7 equities and U.S. Treasury rates. We compare univariate forecasts against multivariate ones — and let you run the real forecasts yourself.
By Sanjiv Das, Tarang Goyal & Mohini Yadav
−60%
Rates forecast error
MV vs UV MAPE (4.97% vs 12.33%)
−16%
Equity forecast error
MV vs UV MAPE (7.06% vs 8.44%)
17
Series, jointly modeled
7 equities + 10 Treasury maturities
Explore
Pick a panel, a series, and a horizon, then scrub through a decade of monthly forecasts. Watch the multivariate and univariate errors diverge over time, and see the win-rate across the whole panel — all computed live from real Chronos-2 outputs.
Coming soonBring your own series
Paste any time series and get a live Chronos-2 forecast on demand. This runs the model at request time on a GPU endpoint — being wired up next.
Method
Forecast each series from its own history alone. One input series in, one forecast out — the classic baseline.
Feed all related series together (e.g. the whole yield curve) and let Chronos-2 attend across them. Same model, richer context.
Findings
Across the full rolling evaluation, feeding Chronos-2 related series together lowers error for every rate and every stock — with the largest gains in interest rates. These are the paper's aggregate numbers.
Mean MAPE across all windows, horizons, and 10 maturities, 2000–2025.
Mean MAPE across all windows, horizons, and 7 equities, 2000–2025.
Forecasting stocks and rates jointly — a combined 17-series panel — slightly degrades accuracy versus modeling each asset class on its own. Cross-domain context adds noise, not signal. MAPE, individual panel vs combined:
Robustness
A fair worry: maybe Chronos-2 already saw these series in pre-training. Two checks argue against it. If the gains were leakage, (1) MV and UV would look alike — they don't, and (2) forecasts before the ~2023 training cutoff would beat those after. Instead, post-2023 forecasts are more accurate, on data the model could not have seen. Suggestive, not conclusive — but the multivariate advantage is not an artifact.