Blog
Field notes from production
What we learn running six billion spans a month for teams shipping agents to real customers.

Research
A practical guide to writing evals that catch real regressions
Most eval suites measure the wrong thing beautifully. A field guide to rubrics that actually move when your product breaks.

Product
Guardrails 2.0: blocking bad output without blocking your users
Inline checks used to mean latency you could feel. The new guardrail runtime runs in under six milliseconds at p95.

Engineering
What 6 billion spans taught us about agent latency
Across every customer we host, the slowest part of an agent is almost never the model. Here is where the time actually goes.

Product
Prompt versioning is a deployment problem, not a document problem
Prompts live in a doc, get pasted into code, and nobody can say which version shipped. Treat them like config and the problem disappears.

Company
How we think about trust in autonomous systems
A short note on the principle behind everything we build: an agent should never be more confident than its evidence.
