No black boxes. Every run, every step.

Everything an agent or workflow does is recorded as a structured trace — inputs, model calls, outputs, tokens, every tool call, guardrail hits, approvals and errors. One screen answers what ran, what it used, how long it took and whether a person checked it.

1explorer for agents & workflows
Step-leveltraces on every run
APIaccess to all run data
Gecko Observability dashboard with representative demo data: token usage, agent and workflow runs, success rate, execution time, charts and recent agent activity
The dashboard

Start every day with “what needs me right now”.

The console opens on a live picture of your AI estate — and the human decisions it’s waiting on.

  • KPI cards — conversations in the last 24 hours, agents, published workflows and recent run errors, each a click away from the filtered runs
  • Activity Pulse — the operational mix of agents, workflows, conversations and crawls, with a daily activity timeline
  • Approvals table — every workflow paused for a person, marked waiting or overdue
Unified run explorer

Agents and workflows, in one operational view.

Agent runs and workflow runs are normalised side by side, so operations, service owners and security all read the same numbers.

Per-run detail

Type, source, status, start, duration, step and tool-call counts, models used, input, output and total tokens, and error details.

Governance as a query

Filter by time range, run type, agent or workflow, status, errors, guardrail triggered and human approval state.

Human review, measured

Review state and review duration on every run, average human review time in the summary — and rejections aren’t counted as errors.

Step-level traces

RAG you can explain to an auditor.

Open any run to see its chronological event trace. For website search, the trace shows the exact search terms, every result with its relevance score, the thresholds applied and a link to the source page.

  • Agent runs — model input, output and token usage, every tool call started, succeeded or failed with its payload, and any guardrail triggered
  • Workflow runs — per-step input, output, timing and errors, with failed steps highlighted on the canvas
  • Cancel in-flight runs, and jump between a run and its conversation in one click
Conversations & review

The human-readable layer — with a review queue built in.

Every conversation, from every channel, is a full transcript with its runs attached. Guardrail hits flag conversations for review automatically, so risk surfaces without anyone reading every chat.

  • Open, Closed and Flagged tabs, with full transcripts and per-message token counts
  • Source metadata — Web, API, Phone, Slack or Teams, with channel and thread identifiers
  • Identity status — verified email and expiry shown on the conversation
  • Review queue — automatic flags from guardrails, manual flags, and who reviewed what and when
  • Staff can step in and message the agent directly from the conversation
Content usage

Prove which content earns its keep.

Knowledge items and every crawled page carry a use count and last-used date. See which pages power answers — and which content nobody has ever needed.

Knowledge & retrieval
Proving value

Build a value case finance can test.

Start with a bounded outcome and baseline today’s service. The platform supplies the operational evidence; you combine it with enquiry volumes, handling time and cost data.

Baseline

Measure today’s volumes, response times and handling effort for one process.

Deploy

Put a governed agent or workflow into production on that process.

Observe

Track runs, success, duration, tokens, tool calls and human review time.

Report

Pull run data through the API into the reporting tools you already trust.

From a student’s chat message to the exact page and relevance score behind the answer — in three clicks.

Make value visible

Know whether your AI is working — not just whether it’s running.

We’ll help you pick the first outcome to measure and show you the evidence the platform produces for it.

  • A walkthrough of the run explorer and traces
  • A measurable first use case
  • The baseline you’ll need to prove it
Book a conversation