01 · Scattered
Three tabs, no story
Prompts in one log, tool calls in another, token counts on a provider dashboard. Nobody can read a run from start to end.
Studio concept · fictional product. The company, product and every name and number in it are invented for this piece. About this piece
Open beta · Tuesday, November 10
Tracewick records every prompt, tool call and handoff your AI agents make in production, flags the step that went wrong, and lets you replay the fix before you ship it.
The problem
Your support agent closed the ticket and replied “Done.” It had also refunded $1,240 without asking anyone. You found out from finance, three weeks later.
01 · Scattered
Prompts in one log, tool calls in another, token counts on a provider dashboard. Nobody can read a run from start to end.
02 · Silent
A run can finish, return a 200 and still do the wrong thing. Nothing alerts on a confident mistake.
03 · Unrepeatable
You change the prompt and hope. There is no way to rerun last week's real tickets and see what it fixed, or broke.
How it works
01
Python and TypeScript SDKs wrap your model and tool calls. Already emitting OpenTelemetry? Point it at Tracewick instead.
# two lines, then every run shows up import tracewick tracewick.init(project="support-agent")
02
Each prompt, tool call and handoff lands on a single timeline, in order, with inputs, outputs, tokens, time and cost.
03
Write rules in plain English. Tracewick checks every run against them and points at the exact step that broke one, even when the run “succeeded”.
Rule · Refunds over $500 need an approval step.
Broken at step 6 · refunds.create · $1,240.00
Product tour
01 · Connect
Install the SDK, name the project, deploy. Every run your agent makes shows up in the list within seconds.
02 · Read the run
Plan, look up, search, decide, act, reply. Open any step to see exactly what went in and what came out.
03 · Find the step
The agent read the refund policy at step 4, then refunded $1,240 at step 6 anyway. Tracewick shows why, with the evidence.
04 · Replay the fix
Replays use the recorded tool responses, so nothing real is refunded. Ship when the diff is green.
Use cases
Support agents
What gets litActions over a limit, promises the policy doesn't allow, loops that reopen the same ticket.
Back-office agents
What gets litAmounts that don't match the PO, the same invoice paid twice, a scanned PDF read upside down.
Coding agents
What gets litTests edited until they pass, runaway retries, a $40 run that should have cost 40 cents.
Research & sales agents
What gets litSources that don't exist, stale numbers, an email drafted for the wrong account.
who get paged when an agent does something strange, and need the run, not a hunch.
running several agents on several model providers, who want one place to see them all.
who have to answer “what exactly did the agent do?” in a sentence, with proof.
Security & data
Traces hold your prompts, your customers' words and your tools' outputs. We treat them that way.
The SDK masks emails, card numbers and API keys on your servers, before anything is sent. Add your own patterns.
Choose where traces are stored and how long they are kept. When they expire, they are deleted, not archived.
Run the collector and the storage in your own cloud account. Only the app talks to us, and it never sees trace contents.
Single sign-on, roles per project, and a log of who opened which trace and when.
Your traces are not used to train any model, ours or anyone else's. Not in beta, not later.
TLS for every connection and encryption at rest for every store, with keys rotated on a schedule.
Tracewick is in beta and has not completed a third-party security audit. We'll publish reports when we have them, not badges before.
Pricing
Builder
$0
For one agent in production.
Team · most teams start here
$400/ month
For teams shipping several agents.
Scale
Custom
For regulated data and big volumes.
Beta pricing. Every plan is free until January 2027. All prices are part of the concept.
FAQ
Each model call (prompt, response, tokens, time, cost), each tool call (inputs and outputs) and each handoff between agents, grouped into one run with a start and an end.
If your agent runs on Python or TypeScript, the SDK wraps model and tool calls directly, whatever the framework. Already sending OpenTelemetry spans? Point them at Tracewick. It works across model providers.
No step waits on Tracewick. The SDK batches and sends in the background. If Tracewick can't be reached, your agent carries on and the SDK retries later.
By default a replay uses the tool responses recorded in the original run, so no real tool is called. You choose which tools, if any, to call live.
Write them in plain English, like “Refunds over $500 need an approval step”, or as code. Tracewick checks every run against them and flags the step that broke one.
No. Tracewick is a studio concept by Artrix Studio, made to show a complete product launch. The company, product and every name, number and screen are invented. The waitlist form doesn't send anything.
Open beta · November 10
Join the waitlist and we'll send your invite on launch morning. One email, no drip campaign.