4K AI engineer - Cyril Nwachukwu

CYRIL NWACHUKWU
- AI ENGINEER

Geist Mono →
#FF5500
02

Telemetry

5-8C15 Industrial telemetry 03
W N E
Geist Mono telemetry
230 mm
300 mm
120 mm
120 mm
+ + +

Project Cartridges

Tactile Hardware Modules // Click to engage Video Oscilloscope

agentpipe

OPEN SOURCE // PYTHON
Ticket LangGraph DAG Fixer Loop CI Tests 237 PASS

Open-source agentic coding pipeline with crash-safe resume, per-role model routing, OpenTelemetry content-hash prompt caching (92% discount), and dual false-pass/false-block eval gates.

apexops

DISTRIBUTED // OUTBOX
Webhook Postgres Outbox Idempotent Billing

Production lead orchestration pipeline with signed webhook ingress, Postgres transactional outbox queue, event sourcing, idempotent billing, and evaluated AI guardrails.

shopify-ai-ops

5 MCP // HMAC-SHA256
MCP 1 MCP 2 MCP 3 MCP 4/5 SQL Engine 0% Hallucination HMAC-SHA256 Safety Gate

Autonomous ecommerce operations agent built with LangGraph v1, 5 Model Context Protocol servers, deterministic SQL analytics to eliminate math hallucinations, and HMAC-SHA256 safety gates.

CoachOS

N8N // GEMINI // SUPABASE
Enquiry Multi-Channel n8n Engine Contracts & Sync Churn Scoring 0 Admin

Complete operational foundation on n8n, Supabase, and Gemini handling enquiry capture, automated contracts, payment sync, session reporting, and churn scoring. Reduced admin from ~15h/wk to near zero.

PropVid

MULTIMODAL // CLOUD RUN
Photos Vision Agent Script Draft ElevenLabs Voiceover MP4

LangGraph multi-agent pipeline for real estate listing videos. Vision agent analyzes property photos, drafting agent writes scripts, ElevenLabs generates voiceovers, and Creatomate renders on Google Cloud Run.

AI Brand Agent

RAG // PINECONE
Docs & Web Chunking pgvector Cosine Index Intent Router Chat & Lead

Production RAG architecture on Pinecone and pgvector with automated document ingestion, intent-based lead capture, and human-in-the-loop escalation over FastAPI.

Sales Intelligence

AUTONOMOUS ENRICHMENT
Scraping Hub Multi-Channel Enrichment Agent Fit Scoring Qualification CRM Push

Autonomous sales intelligence engine that queries multi-source prospect platforms, runs automated company fit evaluations, and delivers verified lead reports directly to sales pipelines.

ANALOG TELEMETRY DECK

Live Calibrated Needles

Continuous telemetry with mechanical needle dials responding to live events and simulation states without re-rendering lags.

Throughput 1,800 t/s
Lat. (ms) 90ms
Cache Hit 92%

A green test proves code runs, not that it is correct.

Telemetry Sandbox

astream_events(v2) Streaming Filter // Live Event Packet Flow

PYTHON // astream_events_filter.py UTF-8
01async for event in graph.astream_events(version="v2"):
02    kind = event["event"]
03    
04    # 1. Lifecycle Events (Update progress bar live)
05    if kind in ("on_chain_start", "on_chain_end"):
06        node = event.get("metadata", {}).get("langgraph_node")
07        ui.set_stepper_status(node, status=kind)
08        
09    # 2. Filter user-facing tokens; suppress background judges
10    elif kind == "on_chat_model_stream":
11        node_name = event["metadata"].get("langgraph_node")
12        if node_name in ("draft", "worst_first_fixer"):
13            chunk = event["data"]["chunk"]
14            ui.render_token_stream(chunk.content)
15            
16    # 3. Deterministic Barrier Synchronization
17    elif kind == "on_eval_barrier":
18        scores = event["data"]["scores"]
19        if scores.weighted_mean() >= 0.75:
20            return Command(goto="open_pull_request")
LIVE OPENTELEMETRY SPAN STREAM SOCKET: ACTIVE
3-COLUMN BROADSHEET EDITORIAL SPECIAL ENGINEERING DISPATCH VOL. IV · NO. 42

The Daily Telemetry

The 3 Invariants of Production AI Agents

It is 2:00 AM. Your agent spent $45 in API credits rewording the same sentence six times. Here is how to survive in production.

Most agent tutorials show you how to build a prototype in twenty lines of code. Almost none show you how to make it survive production traffic. When you build stateful, cyclical workflows in LangGraph, you need three non-negotiable architectural anchors:

Rule 1: Deterministic Control Flow
An LLM is a probabilistic token generator. It should never decide whether a loop terminates or which branch executes. Pure Python checks boolean policies, score thresholds, and iteration ceilings.

Rule 2: Channel Reducers over Shared Variables. Under Pregel execution, parallel nodes do not mutate a shared dictionary in memory. They emit delta updates during a superstep. If two parallel evaluator judges return updates to the same channel without an explicit reducer (Annotated[dict, merge_evaluations]), the last node silently erases the first.

Rule 3: Side-Effect Isolation. External mutations—sending emails, charging cards, calling webhooks—belong in the host wrapper that invokes the graph, never inside node bodies. Otherwise, those side-effects fire repeatedly during test runs, rollbacks, and replays.

Stop Letting The LLM Do The Math

For a while, I had my evaluation agent's LLM judge calculate its own final score. Made sense at first: it already knew the three dimensions. Then came the drift.

For a while, I had my evaluation agent's LLM judge calculate its own final score. It made sense at first: it already knew the three dimensions, so why not have it add them up too? Then I noticed the exact same report, evaluated twice, would sometimes come back with a slightly different final number. Nothing dramatic. Just enough drift to make routing decisions unreliable.

Probabilistic Variance vs. Fixed Weights:
The LLM is probabilistic. Every time you ask it to do something deterministic—arithmetic, applying a fixed weight, enforcing a threshold—you are letting variance into a decision that should never vary.

Moving the Line: The LLM evaluates: it scores each dimension, provides a reason, and emits a confidence signal. Pure Python does everything after that: validates the schema, applies the fixed weights, checks the hard gates, and executes routing. Deterministic, cheap, auditable, testable.

That single boundary—LLM judges, Python decides—has mattered more than any other decision in this build. Where in your pipeline is a probabilistic model quietly doing a job that belongs to plain, reliable code?

The Judge Needs A Check Too

An LLM judge can look excellent during development and still not be production-ready. 10 out of 10 on known examples and 32 out of 50 on unseen ones.

An LLM judge can look excellent during development and still not be production-ready. It might have learned the patterns in the examples you gave it without actually generalizing the underlying skill of evaluation.

You catch this the same way you catch it in the agent itself: test the judge against unseen cases and against trusted reference labels. 10 out of 10 on known examples and 32 out of 50 on unseen ones tells you the judge memorized; it didn't learn.

The Critical Agreement Trap:
Judge A scores 96% overall with 99% agreement on critical cases. Judge B scores 98% overall with only 80% agreement on critical cases. Judge B wins the headline number; Judge A is the one you actually want in production.

When a judge sits on a decision with real consequences, being reliable where it counts matters far more than being right on average. Not all evaluator disagreements cost the same. Weight your judge evaluations accordingly.

A Number And Its Certainty Are Two Different Signals

"Score: 0.84, confidence: 0.95" and "Score: 0.84, confidence: 0.42" are not the same evaluation, even though the score is identical.

The first evaluation says: I evaluated this and I am sure. The second says: I evaluated this and I am not sure I got it right. Treating them identically is one of the most common failure modes in automated eval pipelines.

Track confidence per dimension, not as one blended number for the whole evaluation. Accuracy, clarity, and completeness can each carry a different level of certainty. Lumping them together buries exactly the information you need: which specific part of the judgment is actually shaky.

Actionable Routing vs. Vague Signals:
"Confidence: low" tells you almost nothing. "The transcript doesn't specify whether the client's location requirement is mandatory" tells human reviewers exactly what to look at.

Low confidence on any dimension is an automatic signal to route to a human, along with the specific reason the judge was uncertain. And never take stated confidence at face value: a model claiming 97% confidence needs to be empirically checked against real outcomes before you trust it.

One Number Can Hide A Real Problem

Overall score: 95%. Looks like a healthy system. Break it down: normal cases score 99%, critical cases score 60%.

The 95% headline was hiding a serious issue the entire time, because normal cases outnumber critical ones and drag the average up. This is why a single aggregate score is never enough for production systems.

You need to slice performance by category: critical requests, ambiguous requests, constraint-heavy requests, multi-step requests, long transcripts versus short ones, and cases the system hasn't seen before. Each slice answers a different question about where the system actually struggles.

Averages vs. Slices:
The global average tells you how the system performs on average. The slices tell you how it performs where it matters. Those are not the same question.

Only checking the headline number means you will ship a system that looks fine on a dashboard and fails in exactly the situations you can least afford it to. Granular slicing is the only antidote to false security.

Beyond The Prompt: The 70K Token Blowup

The entire build of agentpipe started from a single number measured on a real pipeline: 70,000 tokens of input to produce 100 tokens of output.

The agent worked fine. It just cost a fortune, quietly, and nobody noticed because nothing ever threw an error. Money was the only symptom. Rebuilding it with measurement-first observability revealed the core fixes:

1. Rebuild, don't accumulate. That 70k blowup happened because every retry re-read all previous attempts. My loop rebuilds its context from the repository as it is now, so attempt four costs about the same as attempt one.

"It works" needs a number:
Putting stable content first earned a caching discount: 92% of repeated prompts billed at the cached rate. But only above the provider's ~1,024-token threshold. The honest version has numbers, not claims.

2. The counter is a hint; the repo is the truth. I made the loop survive a crash by refusing to trust my own attempt counter and always re-checking the actual files on disk. You cannot silently skip work you never assumed was finished.

Passing Your Test Set Proves Almost Nothing

Run your agent against 10 familiar examples and score 100%. Feels like proof. It isn't.

Run your agent against 10 familiar examples and score 100%. Feels like proof. It isn't. That result tells you the system got good at those 10 examples. It says nothing about the cases it hasn't seen, and production is mostly cases it hasn't seen.

A good evaluation set covers normal cases, edge cases, ambiguous requests, constraint-heavy requests, and cases you would consider critical if they went wrong. The question is not "can it pass these examples," it is "does it behave correctly across the kinds of situations it will actually meet?"

Dev Data vs. Holdout Data:
If a prompt change pushes your dev score from 95% to 99% but your holdout score drops from 91% to 82%, that is not progress. That is the system getting better at the examples you are staring at and worse at everything else.

Holdout data is never used for tuning; it is only used to check generalization. Holdout performance is the closer approximation of what production will actually look like.

Why The Weights Aren't 33 / 33 / 33

Once task completion passes, you still need to know how good the report actually is. That's where multi-dimensional scoring comes in.

Once task completion passes, you still need to know how good the report actually is. That's where multi-dimensional scoring comes in: Accuracy (35%), Clarity (35%), and Completeness (30%). Three questions, three deliberate weights.

Accuracy and clarity carry equal weight because both are non-negotiable to the person reading the report. A technically accurate answer they cannot understand has limited value. A clear answer that is wrong is worse.

Deliberate Failure Costing:
Completeness sits lower on purpose. An imperfectly complete report is a smaller problem than a report that is directly wrong or unclear. Decide, on purpose, which failures cost more, instead of defaulting to equal splits because equal feels fair.

The specific numbers matter less than the exercise behind them: decide, on purpose, which failures cost more, and let the mathematical weighting reflect that consequence.

The Production Frame: LLM Evaluates, Python Decides

The difference between a toy demo and an industrial system is not the model you call. It is the deterministic frame you place around it.

When you distill hundreds of hours of debugging stateful agents and evaluating automated workflows, one rule eclipses all others: keep probabilistic models inside narrow, structured boundaries, and let deterministic code govern state transitions.

The LLM is tasked with what it excels at: semantic parsing, extraction, comparative judgment, and structured reasoning with confidence metadata. Pure Python handles everything downstream: DAG channel reducers, SQLite/Postgres checkpointer state, OpenTelemetry span baggage, and fail-fast evaluation gates.

The Core Tenet of agentpipe:
Never trust a model to calculate arithmetic, decide loop terminations, or execute side-effects. Wrap probabilistic intelligence in a crash-safe, deterministic machine.

Software fundamentals matter more in the age of AI than ever before. If you build stateful agent workflows, engineer the deterministic frame first, and the models will finally perform reliably in production.