Writing
Technical notes from the team behind Hotato, local call forensics and regression guards for AI agents: trace, evaluate, test, and gate any AI agent. Each post starts from a public builder report, names the failure class, and ends at a command you run on one recording.
Both sides can talk at once now. What is your voice test suite still measuring?
Full-duplex moved the failures from wording to timing and floor control. A transcript flattens two overlapping voices into one column of words and throws the timing away.
Read the post →The transcript passed. The call failed. A walk from a green log to a red gate.
Rank the moments in one recording, label the catch, and keep it as a contract: the walk from a clean-reading log to an exit 1 that gates the next pull request.
Read the post →Your agent said the refund went through. Only the backend knows if it did.
Say-do grading from trace spans and the system of record, never the model's summary of itself. Includes the authorization and security-state sub-pattern.
Read the post →Every post ends at the same front door, because the loop starts with one call:
hotato autopsy ./call.wav