AgentOps: observing the reasoning, not just the result
Quick answer:
With a model that answers, storing the answer is enough. With an agent that acts, the answer is the last thing that happens and almost never where the failure is.
When an agent returns a wrong result, the useful question is not "what did it say?" but "why did it decide that?". And that question only has an answer if the path was instrumented, not the destination.
What you lose by storing only the answer
An agent resolves a request in several steps: it breaks down the task, picks sources, discards some, calls a tool, evaluates what comes back and decides whether to continue. If the log holds only input and output, all of that is gone.
When the incident arrives — and it will — the team has to reproduce it by hand, with a non-deterministic model that may no longer decide the same way. It is the equivalent of debugging a distributed system with a single print at the end.
What to instrument
Every reasoning step, as a trace
The pattern that works is the distributed trace: the request is a trace and each step a span. For every span, record what was decided, on what input and how long it took. Nothing needs inventing: OpenTelemetry does the job, and the platform team already knows how to read it.
Every tool call
Which tool, with which parameters, what it returned and whether it failed. This is where most real problems show up: the API that changed its contract, the timeout that expires, the parameter the agent fills in wrongly every single time.
Cost, in the same trace
Input and output tokens, model used and estimated cost per step. Sitting next to everything else, it reveals that 80% of the spend comes from one retry nobody had looked at. Split into a separate dashboard, nobody sees it.
The evidence behind the decision
Which sources it consulted and which it used to answer. Without this it is impossible to tell a hallucinating model from a model that answered correctly using data that was wrong. Two different problems, two different fixes, constantly confused. It is the same traceability we demand of any data in a multi-cloud environment, applied to reasoning.
The signals that actually warn you
Four indicators anticipate almost every incident:
- Human intervention rate. If it rises, the agent has started failing at cases it used to resolve.
- Steps per task. An agent needing twice as many steps for the same job is going in circles.
- Cost per transaction. It rises before any other signal when something loops.
- Evidence drift. The agent starts leaning on different sources for the same questions: something changed upstream.
None of the four requires new technology. They require deciding that somebody looks at them.
Circuit breakers: when to stop the agent
Observing without being able to act is of little use. Three breakers worth having from day one:
- A cap on steps and spend per task. If exceeded, the task ends and escalates to a person.
- Credential withdrawal on anomalous behaviour, such as a call volume far above normal. That only works if each agent has its own identity and permissions.
- A business-rule stop: actions that must never execute without human confirmation, however confident the agent claims to be.
The breaker is not a system failure: it is the system working. An agent that never stops is not a reliable agent, it is an unwatched one.
Who looks at all this
The usual organisational mistake is leaving observability to the team that built the agent. That works while there is one. With fifteen, you need a single place where all of them are visible, with the same indicators, and somebody watching with the same discipline applied to production systems.
The same rule applies as to any data product: if it has no owner and no dashboard, it is not in production, it is in extended testing.
How Galde can help
We instrument the agents we build with Palantir AIP with per-step traces, cost per transaction and automatic breakers, and we hand the dashboards to the internal team. We apply the same standard we use to take AI with Palantir AIP to production: nothing counts as finished until somebody can explain why the system decided what it decided.
Conclusion
An agent without observability is not an autonomous agent: it is an unsupervised one, which is not the same thing. The difference between the two is precisely the ability to reconstruct a decision after it happens.
Instrumenting the reasoning costs a few days at the start of the project. Reconstructing it afterwards, without data, costs weeks and sometimes cannot be done at all.
Frequently asked questions
What is AgentOps?
It is the practice of operating AI agents in production: instrumenting their reasoning steps and tool calls, watching cost and drift, and having mechanisms to stop them when they behave anomalously.
How does it differ from traditional observability?
The object under observation is not only infrastructure but the decision itself. Besides latency and errors, you need to record which steps the agent took, which sources it used and which tools it called, because that is where failures usually sit.
Do I need a specific tool?
Not necessarily. Standard distributed tracing plus a metrics dashboard covers most of it. What you do need is to decide what gets recorded at each step, which is a design decision rather than a purchase.
Which signal warns before an incident?
Cost per transaction. It rises before the error rate when an agent loops or starts retrying, which is why it belongs on the same dashboard as everything else.



