Replay a decision
The selected-step lab compares original and edited context across repeated model responses. Repeat rates describe observed behavior; they do not recover private reasoning.
OBSERVABILITY FOR CODING AGENTS
See what your agent saw, what it proposed, and what actually ran. Then change the context and explore what happens next.
“Fix the failing authentication test.”
src/auth.ts · tests/auth.test.ts
Bash
gVisor exec checkpoint · process exit
What did it see? / What did it try? / What actually ran? / What would change?
01 / INSPECT
A final answer only tells part of the story. RayTrace connects prompts, context, tool calls, and runtime evidence so you can inspect the steps that led there.
Inspect the input snapshot for a model request: the prompt, tool definitions, prior results, and file excerpts included in its context.
// Excerpt included in the model request
test('session has not expired', () => {
expect(session.expiresAt)
.toBeGreaterThan(Date.now());
});A file in the repository is not necessarily a file the model saw.
02 / EXPERIMENT
Was a piece of context useful? Select a step, edit or remove what the agent saw, and compare the continuation with the original run.
The selected-step lab compares original and edited context across repeated model responses. Repeat rates describe observed behavior; they do not recover private reasoning.
For eligible sandboxed Claude Code sessions, Playground restores the project before a step and runs a fork. Compare file changes, steps, and an optional check command.
Models can vary even when nothing changes. Playground can interleave edited runs with unchanged runs, so you can examine that variation alongside your experiment.
KNOW THE BOUNDARY Forks need a recorded sandbox session with snapshots. Decision replays and Codex continuations have different setup and provider requirements.
03 / UNDER THE HOOD
A local proxy records model exchanges. SQLite connects agent intent with runtime evidence. The dashboard gives you a way to inspect both.
* OpenRouter is optional experiment routing. Claude Code sandbox sessions also use managed hooks to capture their transcripts.
Trace data stays in a local SQLite database. Content-addressed payloads reuse repeated context instead of storing the same history for every request.
A gVisor sandbox inside a Lima Linux VM independently collects process-start and exit events. The forwarder links them to tool calls using markers or command and timing evidence.
A process starting or exiting successfully does not prove the task was done correctly. Missing evidence means unknown. RayTrace does not reconstruct private chain of thought.
04 / GET STARTED
Start with local request capture. Add sandbox execution when you need independent process evidence.
Built for experimentation. RayTrace is an early prototype for local debugging and evaluation. Captured prompts and file contents can be sensitive; the local services do not yet provide authentication or encryption at rest.
npm installnpm run proxynpm run devRequires Lima and Docker. From your RayTrace checkout:
npm run sandbox:setup
npm run dev:allFor a Claude Code sandbox session:
npm run dev:all -- --claudeOpenRouter-backed decision experiments also require your own API key and provider configuration in the project’s local environment file.