Time-series foundation models for finance

Does giving a forecasting model more series make it more accurate?

A study of Amazon's Chronos-2, a pretrained time-series transformer, on Magnificent-7 equities and U.S. Treasury rates. We compare univariate forecasts against multivariate ones — and let you run the real forecasts yourself.

By Sanjiv Das, Tarang Goyal & Mohini Yadav

−60%

Rates forecast error

MV vs UV MAPE (4.97% vs 12.33%)

−16%

Equity forecast error

MV vs UV MAPE (7.06% vs 8.44%)

17

Series, jointly modeled

7 equities + 10 Treasury maturities

Explore

Test it yourself

Pick a panel, a series, and a horizon, then scrub through a decade of monthly forecasts. Watch the multivariate and univariate errors diverge over time, and see the win-rate across the whole panel — all computed live from real Chronos-2 outputs.

Loading real Chronos-2 forecasts…

Coming soonBring your own series

Paste any time series and get a live Chronos-2 forecast on demand. This runs the model at request time on a GPU endpoint — being wired up next.

Method

Univariate vs multivariate, one foundation model

Univariate

Forecast each series from its own history alone. One input series in, one forecast out — the classic baseline.

Multivariate

Feed all related series together (e.g. the whole yield curve) and let Chronos-2 attend across them. Same model, richer context.

The setup

  • • Input window n = 252 trading days (rolling)
  • • Horizons m = 21 and 63 days (1 & 3 months)
  • • Panels: Mag-7 equities, 10 Treasury maturities, and both combined
  • • Metrics: MAPE (primary) and RMSE on realized values
  • • Zero-shot: Chronos-2 is not fine-tuned on these series

Research questions

  1. 1.Do multivariate (MV) inputs beat univariate (UV) ones when a foundation model forecasts both?
  2. 2.Is the MV advantage larger for stocks or for interest rates?
  3. 3.Does forecasting stocks and rates jointly — a small “world” model — help or hurt?
  4. 4.Are the gains an artifact of pre-training leakage, or real?

Findings

Multivariate inputs consistently beat univariate ones

Across the full rolling evaluation, feeding Chronos-2 related series together lowers error for every rate and every stock — with the largest gains in interest rates. These are the paper's aggregate numbers.

Rates

−60% MAPE
Multivariate4.97%
Univariate12.33%

Mean MAPE across all windows, horizons, and 10 maturities, 2000–2025.

Stocks

−16% MAPE
Multivariate7.06%
Univariate8.44%

Mean MAPE across all windows, horizons, and 7 equities, 2000–2025.

Treasury rates — MAPE by maturity

3-Month
23.55%
31.86%
6-Month
8.33%
17.70%
1-Year
5.72%
13.26%
2-Year
3.34%
12.70%
3-Year
2.43%
11.81%
5-Year
1.71%
9.72%
7-Year
1.34%
8.27%
10-Year
1.14%
7.03%
20-Year
0.99%
5.72%
30-Year
1.13%
5.20%

Equities — MAPE by ticker

Apple
5.99%
7.20%
Amazon
5.67%
7.07%
Alphabet
4.78%
6.35%
Microsoft
3.68%
4.66%
Netflix
10.40%
11.45%
NVIDIA
9.06%
11.07%
Tesla
10.98%
12.42%
Multivariate Univariate(lower is better)

But a bigger “world” model doesn't help

Forecasting stocks and rates jointly — a combined 17-series panel — slightly degrades accuracy versus modeling each asset class on its own. Cross-domain context adds noise, not signal. MAPE, individual panel vs combined:

3-Month23.04% → 23.92%
2-Year3.66% → 3.92%
10-Year1.28% → 1.26%
NVDA7.43% → 7.61%
MSFT3.05% → 3.18%
TSLA10.90% → 11.14%

Robustness

Is this just training leakage?

A fair worry: maybe Chronos-2 already saw these series in pre-training. Two checks argue against it. If the gains were leakage, (1) MV and UV would look alike — they don't, and (2) forecasts before the ~2023 training cutoff would beat those after. Instead, post-2023 forecasts are more accurate, on data the model could not have seen. Suggestive, not conclusive — but the multivariate advantage is not an artifact.