Skip to content

Engineering notes

Keelo publishes method, not customers.

Keelo does not publish customer names. It publishes method. These are working notes on the engineering that keeps a fleet of production agents accurate, observable and repairable — the arithmetic behind why agents that demo well degrade in production, the failure classes worth instrumenting for, and what a machine-written fix has to clear before it reaches anything real.

  1. 25 July 2026

    7 min read

    Why an eight-stage agent pipeline at 95% per stage returns 66%

    Stage accuracy multiplies. Almost every hard engineering decision in a production agent follows from that one fact, and most agent architectures are designed as though it were not true.

    • Reliability
    • Evals
    • Agent architecture
  2. 25 July 2026

    6 min read

    The eleven ways a production agent fails

    A named failure taxonomy is the difference between an agent you monitor and an agent you can repair. These are the eleven classes Keelo's watchdog detects, and what each one actually catches.

    • Observability
    • Reliability
    • Production agents
  3. 25 July 2026

    6 min read

    What a machine-written fix must clear before it reaches production

    An agent that writes its own fixes raises exactly one question worth asking. The answer should be enforced by code paths, not promised by policy — here are the four gates that enforce it.

    • Autonomy
    • Safety
    • Self-healing systems
  4. 25 July 2026

    5 min read

    Approval is the highest-signal event in an agent system, and it is free

    Deterministic checks produce proofs. Model judges produce opinions. Neither knows whether the person who asked was actually served — which is why the approval a reviewer already gives should be the primary training signal.

    • Evals
    • Human-in-the-loop
    • Feedback loops

These notes describe how Keelo builds. They are written from production systems running in customer environments, and they identify no customer, no vendor, and no deployment. Where a note gives a number, that number is either arithmetic you can reproduce or a threshold Keelo enforces — never a measured customer outcome.

Get started

Bring the workflow that matters too much to leave as a prompt experiment.

The first conversation is about which workflow is worth encoding, what has to stay human, what has to be governed, and what system needs to exist around the model. It is a technical conversation with the person who will build the thing.

Keelo takes on a small number of deployments. The work has to matter — to the business, to the people doing it, and to Keelo.

Direct: Edward@keelo.ai