RayTrace

RAYTRACE DOCS

From first capture
to a better question.

Set up local capture, inspect a coding-agent trace, and explore how decisions change when you change the context.

RayTrace runs locally. This website explains the project; the dashboard and proxy run on your machine.

What RayTrace records

Model requests and responses, input context, proposed tool calls, and reported results. Optional sandbox instrumentation adds independently collected process-start and exit events. These are separate sources of evidence.

Quick start

You need Node.js 22.13 or newer, npm, a RayTrace checkout, and a compatible client that can send model requests through the local proxy.

1. Install and start the proxy

cd raytace
npm install
npm run proxy

The default proxy address is http://127.0.0.1:8797. Configure your compatible client’s base URL to use it. Some clients expect the /v1 suffix.

2. Open the dashboard

In a second terminal, from the same checkout:

npm run dev

Open http://localhost:3000 and send a request through the proxy. The database is created automatically at .raytace/evidence.db.

Capture & inspect

The current dashboard connects six views: Prompts, Tool calls, Sandbox, Context, Playground, and a reasoning view. Start with a prompt, then inspect its proposed actions and the context supplied for each model request.

  • Prompts: browse captured prompts and responses.
  • Tool calls: inspect tool names, arguments, reported results, and attribution.
  • Context: see the input available at a selected step.
  • Sandbox: inspect the independently collected process timeline.

Visible reasoning or generated explanations do not reconstruct the model’s private chain of thought.

Sandbox execution

The optional runtime uses Lima, Docker, and gVisor. Set up the dedicated Linux VM once, then start the local services and sandboxed agent:

npm run sandbox:setup
npm run dev:all

To start Claude Code instead of Codex:

npm run dev:all -- --claude

Claude Code sandbox sessions use a managed hook to tag Bash commands with tool-call IDs and capture session transcripts. Only commands in the instrumented sandbox are monitored.

Service controls

npm run dev:all -- --status
npm run dev:all -- --no-codex
npm run dev:all -- --stop

--stop also frees ports 8797, 8799, and 3000, including processes holding those ports that the script did not start. Export sandbox work deliberately; it does not automatically sync to the host.

Experiments

Decision replay

The selected-step lab edits or removes context and compares repeated model responses against the original context. It supports 2–20 runs per version. Matching uses exact tool names and arguments or exact answer text; repeat rates are observations, not causal probabilities.

Claude Code Playground

For a captured sandbox session with snapshots, select a model step, change earlier tool results or project files, and continue in a separate fork restored from before that step. Compare tool calls, answers, file diffs, and an optional check command.

Run repeated forks and unchanged baselines to examine normal model variation. A check result reflects the configured command’s exit code, so choose a check that measures the outcome you care about.

Codex execution continuation

The lab’s Codex continuation is a different workflow. It reconstructs an edited transcript in an isolated copy of current Git-listed files, rather than restoring an exact historical session. It requires local Codex and OpenRouter mode; workspace changes stay in the run directory.

Model calls for summaries and experiments can incur provider charges. A changed result is evidence for investigation, not proof that a single context item caused it.

Configuration

Use the project’s .env.example as the configuration reference. Never commit your API keys.

cp .env.example .env
# Set your own OPENROUTER_API_KEY in .env
# Enable RAYTACE_PROVIDER=openrouter for experiment routing
npm run proxy
VariablePurpose
RAYTACE_PORTProxy port; defaults to 8797.
RAYTACE_DBOverride the local SQLite database path.
RAYTACE_PROVIDERSet to openrouter for optional experiment routing.
OPENROUTER_API_KEYYour key for OpenRouter model calls.
RAYTACE_OPENROUTER_MODELConfigured experiment model or alias.
RAYTACE_OPENAI_UPSTREAMOptional OpenAI-compatible upstream.
RAYTACE_ANTHROPIC_UPSTREAMOptional Anthropic upstream.

Existing shell environment variables take precedence over the environment file. Restart the proxy and start a fresh agent session when changing providers.

Evidence & privacy

A proposed command, an agent-reported result, a process-start event, and a verified task outcome are different facts. gVisor can witness execution, but it does not establish that the agent solved the task.

  • Missing or unattributed execution evidence means unknown.
  • Shell builtins do not create new executable processes.
  • The prototype does not capture every file operation, network request, or terminal output.
  • Local prompts, arguments, file contents, and outputs may contain sensitive material.
  • The prototype does not provide authentication, encryption at rest, a retention policy, or an external immutable audit service.

Credential-shaped JSON fields are redacted before storage, but redaction is not a guarantee that all secrets are removed. Keep the viewer and proxy as trusted local development services.

Troubleshooting

The dashboard has no captures

Check that the proxy is running and your client uses the proxy base URL. An already-running desktop agent is not automatically reconfigured by a new CLI launcher.

A tool call has no sandbox evidence

Confirm that the command ran inside the gVisor sandbox and that the manager, evidence viewer, and proxy are running. Unmatched processes can remain unattributed; this is not proof that a call never ran.

A step cannot be forked

Playground needs a recorded sandbox session and a usable snapshot before the selected step. Ordinary proxy-only captures are not enough.

Experiment routing fails

Check your local API key and model configuration, then restart with a fresh session. Provider or key errors do not automatically fall back to native routing.