Portfolio

My personal dark factory

Agents write the code. I own the spec and the gates. I'm a senior backend engineer and open source author. Below is what it has built so far: a stack of three open source Go projects, from agents to durable workflows to fleet orchestration.

About this factory

A dark factory is a pipeline where a spec goes in and tested software comes out, with agents doing the planning, writing, testing and review in between. Mine is not fully dark, on purpose: the lights stay on where judgment lives. What to build, which behaviors must never break, and what counts as done. Everything between those points is delegated, and delegated hard.

I still think of myself as a craftsman, and I have fully embraced AI-first. The two fit because the craft moves into the specs and the checks. It is also why AI-first is a way of working and AI-only is a way of failing slowly (more in this post).

Spec-driven

A CLAUDE.md per repo as standing orders, plus plans, event and protocol specs, and invariant tables.

Gates, not vibes

Black-box and chaos tests, race detector, linter, and architecture rules enforced by tests.

Craft stays

Design before code, named failure modes, and an append-only backlog for every doubt.

What the factory built

Three open source Go projects, each with one job, layered so that each can be used alone. All three were built with the process above, and each repository carries the specs and gates that made that possible.

Layer 1 · Agents

phero

The chemical language of AI agents.

A Go framework for multi-agent systems. Small composable packages instead of a monolith, with a provider-agnostic LLM layer for OpenAI-compatible endpoints and Anthropic.

  • Agent orchestration with tool execution, chat loops and runtime handoffs
  • Function tools with automatic JSON Schema, skills from SKILL.md files, and MCP servers as tools
  • Memory, embeddings, RAG and vector stores (Qdrant, PostgreSQL/pgvector, Weaviate)
  • A2A and a NATS agent protocol, so agents can be discovered and called over pub/sub
  • LLM middleware (retry, rate limit, guardrails, semantic cache) and typed tracing with OpenTelemetry

How it is specified: package-by-package conventions in CLAUDE.md, end-to-end tests against real services, and an append-only review backlog where each finding is closed with a regression test.

Layer 2 · Durability

packtrail

A durable workflow engine built only on NATS JetStream.

Workflows as data, not as code. You declare a graph and the engine interprets it. Every execution is an ordered event log, and state, timers, jobs and the visibility index are all derived from it.

  • Event-sourced, so time travel, forks, replays and audit come for free
  • Tasks, choices, fan-out and join, human-in-the-loop awaits, dynamic maps, subflows and failure routing
  • Workers in any language over a small versioned JSON protocol
  • Agnostic by design: no agents, LLMs or tokens in the core
  • No database and no coordinator: single-writer appends with expected-sequence checks, and everything else in NATS

How it is specified: an event-log spec, a worker protocol spec, and an invariants table where every behavior points at the test that proves it. Acceptance and chaos suites run against an embedded NATS server, never mocks. Allowed dependencies are pinned by a test.

Layer 3 · Orchestration

stiggy

Multi-agent architectures, deployed on NATS.

Glues phero's agents to packtrail's durable flows. You describe a fleet once, as a YAML file or in Go, and it compiles into flows whose steps run agents. It covers what LangGraph and CrewAI provide, on a substrate that survives crashes and scales out.

  • Models, tools, knowledge and agents with CrewAI-style role, goal and backstory
  • Crews and patterns: supervisor, swarm, evaluator-optimizer, debate and plan-execute, expanded into plain flows
  • Validation that reports every problem with its line, and a compiler you can inspect
  • Execution control from the CLI: signal, resume, fork, rerun, cancel and watch progress

How it is specified: a reference doc and a feature-by-feature coverage map against LangGraph and CrewAI, golden tests for the compiler, and a hard architecture rule: stiggy never touches NATS directly, and a type-based guard test enforces it.