Blog
hedwigaihedwigaihedwigai

Time machine: leak-free RL environments, graded by what actually happened

Time machine is a feature for hedwigai: pick a date, and every source your research reads answers as it would have on that day. Filings not yet filed are hidden. Prices are the ones printed that day. Nothing that came later gets in.

10-K$121.8bn

Two things it is for

Primary research, as of a date

Ask a question the way it could have been answered on a past day, from the filings, prices and records that existed then.

  • What could a buyer have known about a company the day the deal was signed?
  • Would last year's investment memo still hold, written only from what was public then?
  • Re-run a past research brief and see what it missed, and what had not happened yet.
RL environments for forecasting

Set the clock back, let an AI agent place bets, then turn the clock forward and pay out each bet as its answer is published, ending on a result card.

  • Thousands of practice questions a year, every one graded by a record, not a person.
  • The same agent scored with the clock set and unset, to prove nothing leaks.
  • A forecaster trained on one year and tested on the next.

The second needs a word of background, because it is where the industry's attention is.

A grader you don't have to write

Every AI lab now wants reinforcement-learning environments: places where a model practises a task, tries things, and is scored on how it did, thousands of times over. The hard part is never the practice. It is the score: somebody has to know the right answer for every attempt, without being fooled.

An environment needs three things: tasks, a way for the model to act, and a score it cannot cheat. Code is the favourite because tests pass or fail. Most real work has no tests, so environment builders pay experts to grade, or ask another model to, and both can be fooled.

A forecast is different. “Will this company's next annual report show X?” has an answer that arrives on its own, from the world, with a date on it. And forecasting is no toy: a team at Berkeley got a search-and-reason system to the level of a human forecasting crowd, and a team at Lightning Rod Labs trained a 14-billion-parameter model on nothing but how Polymarket questions turned out, and it matched far larger models (sources at the end).

The trouble is that a real forecast takes months to score. So everyone backtests: pretend it is the past, ask about what happened next, score at once. That is exactly what a time machine does, and it only works if the past stays the past.

The future leaks

Traders have a name for a backtest that can see the future: look-ahead bias, and it makes every strategy look brilliant. AI forecasting has the same problem in new forms. Researchers at ETH Zürich's SPY Lab showed that search engines' “before this date” filters let later pages through, and that models know things from after their stated training cutoff. A Northwestern study this August found the standard check for that kind of leak flagging four of five models on questions they could not have seen.

The careful answer so far has been to freeze a news archive and search only that. It works, but news is only part of what an agent reads. It also looks up filings, prices, trial records and drug approvals, and each one is a door the future can come back through.

Set the clock on the shell, not the search engine

hedwigai ask does all of its looking up through one shell: a locked-down command line where each data source is a program, such as biotech for SEC filings, FDA approvals and clinical trials. One door is the point. If every source comes through the shell, one clock on the shell covers every source.

The everyday shell knows only now. Its one clock function is time, and it reads the real date. The clock lives in a separate place, the gym: a shell session opened at a past date, where time says that date and two programs exist that exist nowhere else. bet records a probability that something will happen, and fast-forward turns the clock on, resolving every bet whose answer has been published by then. Every other program follows one rule there: answer only from what had been published by the gym's date. SEC filings make it easy, because every one carries the day it was filed. Add a program for share prices, as printed on the day, and an agent can reason about a company's value the way an investor could have then.

# a gym session: the shell opens with its clock at 2025-10-10
$ time
2025-10-10

$ biotech company marketcap "Vertex Pharmaceuticals"
public_float $121.8bn  as of 2024-06-28  (10-K filed 2025-02-13)

$ bet "Vertex's next 10-K shows a public float above $120bn" 0.80
bet placed; resolves when the answer is published

$ fast-forward 2026-02-13
resolved: no ($113.4bn, 10-K filed 2026-02-13)  reward 0.36
A sketch of a gym session. Outside the gym, the same lookup returns the $113.4bn filed 2026-02-13, and bet and fast-forward do not exist.

One question, scored

Set the clock to 2025-10-10 and ask: will Vertex Pharmaceuticals' next annual report show a public float (the market value of the shares the public holds) above $120bn? With the clock set, the agent sees 3 years of it, climbing from $71.8bn to $121.8bn. The trend says yes.

Here is what the two annual reports actually say, word for word.

The report filed on 2026-02-13 said $113.4bn. The answer is no. An agent that followed the line loses points; one that hedged keeps more. Nobody had to grade it: the answer is a number in a filing.

Vertex Pharmaceuticals · Form 10-K, fiscal 2024Filed 13 February 2025
“The aggregate market value of the registrant’s common stock held by non-affiliates of the registrant based on the closing price on June 28, 2024 (the last business day of the registrant’s most recently completed second fiscal quarter of 2024) was $121.8 billion.”
Visible with the clock at 2025-10-10Read it on SEC.gov
Vertex Pharmaceuticals · Form 10-K, fiscal 2025Filed 13 February 2026
“The aggregate market value of the registrant’s common stock held by non-affiliates of the registrant based on the closing price on June 30, 2025 (the last business day of the registrant’s most recently completed second fiscal quarter of 2025) was $113.4 billion.”
Hidden at 2025-10-10: this is the answerRead it on SEC.gov
2022filed 2023-02-10$71.8bn
2023filed 2024-02-15$90.7bn
2024filed 2025-02-13$121.8bn
2025filed 2026-02-13$113.4bn ?
the agent can see itfiled after the clock: the answerthe $120bn line
if it had been yes
0.96
if it is no (it was)
0.36

Score is 1 minus the squared miss, from 0 to 1. Saying 50% always scores 0.75; only a probability that leans the right way scores more, and a confident wrong one scores least.

A shell session is an episode

In reinforcement learning, an episode is one go at the task, from start to score. With a time machine, it is simply a shell session. The agent enters a shell whose clock reads 2025-10-10, does whatever research it likes, places its bets, and leaves. Then the clock turns forward, one filing at a time. Each time a bet's answer is published, the bet resolves and pays out. When the last one resolves, the session closes on a result card.

Try it below as the agent: 4 real biotechs, each asking whether the next annual report shows a higher public float than the last one filed before the clock. The answers are what the companies actually filed.

  1. 1Enter

    A gym shell opens at a past date. Every program in it answers as of that day.

  2. 2Work

    The agent researches as it likes: filings, prices, trials, the news of the day. Then it places its bets with bet.

  3. 3Leave

    fast-forward turns the clock on. Each bet pays out as its answer is published, and the session ends on a result card.

clock: 2025-10-10collected 0.00 of 4
Regeneronlast $113.8bnresolves 02-04
Amgenlast $167.6bnresolves 02-13
Vertexlast $121.8bnresolves 02-13
Gileadlast $60bnresolves 02-24

Vertex's float had climbed three years running before the clock, the kind of line a trend-follower bets on. Betting up on all 4, it scores 2.44, below the coin's 3.00. Telling those two apart is what an environment like this is for.

We ran the gym on EDGAR

Before building anything, we simulated the design with the figures EDGAR already publishes: every US filer's public float for five years. One episode per year, with the clock at 1 April and one bet per company: will its next float be higher than its last? That made 15,277 bets over 4 years, every one graded by a later filing.

The RL players learned from payouts alone: bet, fast-forward, collect. Each was trained on three years and scored on the fourth, which it never saw, for each year in turn, and five times over. Scores are the average payout per bet, from 0 to 1; the coin always scores 0.75.

1RL player, late clockclock 1 October, after the answer is measured0.762
2A coin50% on everything0.750
3RL playerits own filings only0.747
4RL player + indexand the S&P 500 up to the clock0.726
5Momentum80% if it rose last year, else 20%0.708
6Trend-follower70% up on everything0.685
Bars start at 0.65; the line is the coin. Red: a clock that leaks.
Held-out yearBetsWent upS&P 500 to 1 AprilRL playerLate clock
20223,98727%+6%0.7520.799
20233,97744%+9%0.7460.757
20243,74350%+18%0.7470.748
20253,57054%+3%0.7410.745
  • The coin is hard to beat. With only what a company's own filings showed, the RL player learned to stay close to 50% and scored 0.747, just under the coin. The rules a person might pick did worse: always lean up, or follow last year.
  • Thousands of bets, 4 draws. The share of floats that went up swung from 27% to 54%, set mostly by the market. Giving the player the S&P 500 up to the clock made it worse (0.726): three years teach it a rule about the market, and the fourth breaks it. A gym needs many clocks, not only many bets.
  • The leak paid best. Move the clock to 1 October and nothing published yet gives the answer, but the floats were measured on 30 June and the index by then already says most of it. The same player scored 0.762, the best of all, and 0.799 in 2022, when the S&P 500 had fallen 17% since the last measurement. It looked like skill and was a leak; it is the first rule below.

What a clock alone doesn't stop

Hiding later documents is the easy part. 5 leaks get past it, and each needs its own rule in the programs:

The answer is fixed before it's published

A float is measured on 30 June and published months later. In between, nothing filed says it, but the day's prices already do.

Rule: Set the clock before the answer is measured, not just before it is published.

The model remembers

A model trained on text from after the clock may simply know the answer.

Rule: Set the clock after the model's training ended, and keep the questions that resolve after that too.

Filings get corrected

A company can amend a filing months later. Today's database shows the corrected figure under the old date.

Rule: Answer with each filing as it was first filed, by its filing date, not the period it covers.

Prices get rewritten

Price histories are adjusted for later stock splits and dividends, so an old price you look up today carries the future in it.

Rule: Serve the price as it was printed that day, unadjusted.

The losers vanish

Lists of companies built today leave out the ones that went bust or were bought. A bet on survival looks safer than it was.

Rule: Build every list as of the clock, delisted companies included.

What comes next

  • The next model: the RL player above is a few weighted numbers. Next is a small decision model, of the kind that makes the yes-or-no calls inside ask, trained the same way on its result cards: it can read what a filing says, not only its figures.
  • The leak check: the same agent with the clock set and unset, on questions whose answers it should not know. If it scores better with the clock unset, the clock is working; if it scores the same, something is leaking.

To try ask, add hedwigai to your AI assistant: mcp.hedwigai.com.

Sources

  • Approaching Human-Level Forecasting with Language Models — Halawi, Zhang, Chen and Steinhardt (NeurIPS 2024): a search-and-reason system scored a Brier of 0.240 against the human crowd's 0.247, on questions asked after its models' training ended.
  • Outcome-based Reinforcement Learning to Predict the Future — Turtel, Franklin, Skotheim, Hewitt and Schoenegger (TMLR 2025): a 14-billion-parameter model trained only on how Polymarket questions resolved matched frontier models on accuracy and beat them on calibration, over 3,300 later questions.
  • Scaling Open-Ended Reasoning to Predict the Future — Chandak, Goel, Prabhu, Hardt and Geiping (2025): questions made from daily news, and an offline news archive for both making them and searching, so nothing after the date leaks in.
  • LLM Forecasting Evaluations Need Fixing — Paleka, Goel, Geiping and Tramèr (June 2025): search engines' date filters let later pages through, and models know things past their stated training cutoff.
  • Temporal Leakage in LLM Backtesting — Zhang and Stadie (August 2026): the usual before-and-after-the-cutoff leakage check failed for four of five models on questions they could not have seen, because it measures how recent a question is, not leakage.
  • verifiers and the Environments Hub — Prime Intellect: an open library and hub where an environment is a set of tasks, a harness and a grader, usable for training and evaluation alike.
  • Who Will Win the RL Environment Market — and Why — Chris Zeoli, Wing VC (January 2026): about twenty startups selling environments to labs, expected to narrow to a few; the moat is stable, repeatable environments and verification that holds up.
  • The simulation reads EDGAR's XBRL frames for EntityPublicFloat, mid-2021 to mid-2025, every filer with a float in two consecutive years, and the S&P 500 closes from FRED. The RL player is a policy over a handful of numbers (the last two changes in its float, its size, and, where given, the index), trained by REINFORCE on the payout alone.
  • Vertex's public float is the EntityPublicFloat figure in each 10-K, from SEC EDGAR (CIK 875320), with each filing's filing date. ask returns the latest of them through biotech company marketcap.