RayTrace

OBSERVABILITY FOR CODING AGENTS

See what your coding agent did.
Understand what changed.

See what your agent saw, what it proposed, and what actually ran. Then change the context and explore what happens next.

Local-firstOpen sourceMIT licensed
raytrace / session inspectorEXAMPLE
SESSION / 0013 events

“Fix the failing authentication test.”

01
Context received INPUT

src/auth.ts · tests/auth.test.ts

expect(session.expiresAt)
  .toBeGreaterThan(Date.now());
02
Tool call proposed INTENT

Bash

$ npm test -- auth
03
Process observed EVIDENCE

gVisor exec checkpoint · process exit

npm test -- authexit 0
[ ] A proposal and an observed process are different facts.
THE QUESTIONS BEHIND EVERY TRACE

What did it see? / What did it try? / What actually ran? / What would change?

01 / INSPECT

A trace you can
take apart.

A final answer only tells part of the story. RayTrace connects prompts, context, tool calls, and runtime evidence so you can inspect the steps that led there.

TRACE ANATOMY

Illustrative example.
No live agent is connected.

MODEL INPUTauth-fix / step 01

Start with what the model saw.

Inspect the input snapshot for a model request: the prompt, tool definitions, prior results, and file excerpts included in its context.

context / tests/auth.test.ts
// Excerpt included in the model request
test('session has not expired', () => {
  expect(session.expiresAt)
    .toBeGreaterThan(Date.now());
});

A file in the repository is not necessarily a file the model saw.

02 / EXPERIMENT

Change one thing.
Follow a new branch.

Was a piece of context useful? Select a step, edit or remove what the agent saw, and compare the continuation with the original run.

CAPTURED STEPOriginal contextPrompt + tool results + files
BASELINEKeep context unchangedRepeat the same step
FORKEdit context or project filesContinue in a separate sandbox
COMPAREWhat changed?Tool calls · answers · file diffs · checks
A

Replay a decision

The selected-step lab compares original and edited context across repeated model responses. Repeat rates describe observed behavior; they do not recover private reasoning.

B

Continue from a snapshot

For eligible sandboxed Claude Code sessions, Playground restores the project before a step and runs a fork. Compare file changes, steps, and an optional check command.

C

Compare against a baseline

Models can vary even when nothing changes. Playground can interleave edited runs with unchanged runs, so you can examine that variation alongside your experiment.

KNOW THE BOUNDARY Forks need a recorded sandbox session with snapshots. Decision replays and Codex continuations have different setup and provider requirements.

03 / UNDER THE HOOD

Local capture.
Connected evidence.

A local proxy records model exchanges. SQLite connects agent intent with runtime evidence. The dashboard gives you a way to inspect both.

>_
Agent / clientOpenAI-compatible traffic
request / response
⌁
RayTrace proxy127.0.0.1:8797
LOCAL
Routes to OpenAIAnthropicOpenRouter*
captured exchanges
▤
SQLite evidence store.raytace/evidence.db
inspect / experiment
◫
Dashboard + experiment enginelocalhost:3000

* OpenRouter is optional experiment routing. Claude Code sandbox sessions also use managed hooks to capture their transcripts.

STORAGE

Keep the history close.

Trace data stays in a local SQLite database. Content-addressed payloads reuse repeated context instead of storing the same history for every request.

OPTIONAL RUNTIME

Observe beyond the transcript.

A gVisor sandbox inside a Lima Linux VM independently collects process-start and exit events. The forwarder links them to tool calls using markers or command and timing evidence.

EVIDENCE BOUNDARY

Observed execution ≠ correctness.

A process starting or exiting successfully does not prove the task was done correctly. Missing evidence means unknown. RayTrace does not reconstruct private chain of thought.

04 / GET STARTED

Your agent.
Your machine.
Your trace.

Start with local request capture. Add sandbox execution when you need independent process evidence.

REQUIRESNode.js 22.13+ · npm · compatible client

Built for experimentation. RayTrace is an early prototype for local debugging and evaluation. Captured prompts and file contents can be sensitive; the local services do not yet provide authentication or encryption at rest.

LOCAL SETUP
  1. Install in your RayTrace checkoutnpm install
  2. Start the capture proxynpm run proxy
  3. Start the dashboard in another terminalnpm run dev
Add the optional gVisor sandbox

Requires Lima and Docker. From your RayTrace checkout:

npm run sandbox:setup
npm run dev:all

For a Claude Code sandbox session:

npm run dev:all -- --claude

OpenRouter-backed decision experiments also require your own API key and provider configuration in the project’s local environment file.

BUILD WITH US

Bring a real agent workflow.
Help shape RayTrace.

Work with the early prototype on the debugging and evaluation questions that matter to your team.

Design Partner