Benten

Voice AI isn't judged by what it says.
It's judged by how it sounds.

Benten is open-source observability for Voice AI. We listen to the actual conversation — the latency, the silence, the interruptions — so you hear what your customers hear.

MIT licensed · Self-hosted · No hosted service, no lock-in

✓ Naturalgap · 480ms
  • HumanHey, can you check my order status?
  • AgentSure — one moment while I pull that up.
Silence after human stops480ms
✗ Abandonedgap · 2.6s
  • HumanHey, can you check my order status?
  • HumanHello?… Hello?
  • SystemCall abandoned
Silence after human stops2600ms
Story 01 — Timing

Two identical sentences.
One feels human. One doesn't.

Two conversations can say the exact same words. The difference isn't language — it's rhythm. Benten measures the rhythm.

Instant

100ms

Feels like a person.

Conversational

400ms

You barely notice it.

Slow

1200ms

Something feels off.

Broken

2500ms

"Hello?…"

Story 02 — The transcript lies

Your transcript says everything went perfectly.
The customer remembers the silence.

Your customers never read the transcript. They hear the pauses, the interruptions, the dead air. Words are not the product. Timing is.

What the transcript shows

User: I'd like to change my flight.

Agent: Sure, I can help with that.

Looks perfect. Ship it.

What Benten hears

User: I'd like to change my flight.

2.4s dead air

Agent: Sure, I can help with that.

The customer already thinks the call disconnected.

Story 03 — Independence

Don't trust your provider.
Measure the conversation yourself.

Every voice provider grades their own homework. Benten is an independent layer of truth above ElevenLabs, Vapi, and Retell — so you can compare providers fairly, and see what actually shipped.

Your provider dashboard

API uptime100%
Requests12,481
Errors0.02%

Says: "Everything is healthy."

What Benten measured

Avg turn latency2,140ms
Dead air rate8.4%
Callers who said 'Hello?'271

Reality: Your customers are hanging up.

Story 04 — Every millisecond

Your Voice AI lives or dies
before it finishes the sentence.

Every delay changes perception. 100ms feels instant. 400ms feels human. 1,200ms feels slow. 2,500ms and they think the line dropped.

100ms
Instant
400ms
Conversational
1200ms
Slow
2500ms
"Hello?"
Story 05 — The missing debugger

Voice has no observability.
Until now.

Cloud providers monitor servers. Voice providers monitor APIs. Neither monitors conversations — and conversations are what your users actually experience.

Software engineershave logs.
Backend engineershave traces.
Frontend engineershave DevTools.
Voice engineerslisten to recordings, hoping to find the problem.
Story 06 — The invisible bugs

Every pause tells a story.
Not because they're metrics.
Because they're the reason conversations feel human.

Your QA team never notices these — they're not crashes, they're half-second hesitations. The bugs your customers notice first are the ones your logs never record.

Dead air

The silence that makes callers think you dropped.

Interruptions

When the model cuts the human off mid-word.

Barge-ins

When the human has to talk over the assistant to be heard.

Latency

The gap from human-stop to agent-start, in milliseconds.

Noise

Line quality, distortion, and background artifacts.

Speech overlap

The two-people-talking-at-once moments.

Story 07 — One conversation

Every provider tells a different story.
The conversation only happened once.

ElevenLabs says one thing. Vapi shows another. Retell exposes something else. Benten sits above all of them and gives you one timeline, one set of metrics, one truth.

ElevenLabs

Their version

Vendor-graded metrics

Vapi

Their version

Vendor-graded metrics

Retell

Their version

Vendor-graded metrics

Benten

One conversation. One timeline. One truth.

Story 08 — The fingerprint

Every conversation leaves
a fingerprint.

Some feel effortless. Some feel awkward. Some make people hang up. Benten shows you exactly why — the shape of every call, the moment it went wrong, the millisecond it fell apart.

"That felt weird" has a cause.
Too much silence. Interruptions. Talking over people. Bad pacing.
Benten measures the moments that make conversations feel human.

MIT · Self-hosted only

No hosted service.
No lock-in.
Just the code.

Benten is 100% open source. Clone it, run it on your own infrastructure, and own your voice observability stack end-to-end. We don't run a cloud. We don't want your call audio.

~/benten
$ git clone github.com/ayusrjn/benten
$ cd benten
$ docker compose up

→ syncing agents from ElevenLabs...
→ syncing agents from Vapi...
→ syncing agents from Retell...

 34 agents discovered
 listening to conversations
 dashboard live at localhost:3000