Toolkit
The toolkit behind the systems
The operation runs on 31 purpose-built skills and agents — each a small, single-job tool rather than one large assistant. Some are autonomous agents (the diagnostic pipeline orchestrates a panel of specialist advisors); most are commands I reach for by hand. They're grouped here by what they do. The longer write-ups in Systems explain how they fit together; this is the inventory underneath them.
Session & governance
The operating rhythm — how every session opens, captures, and closes.
Session governor
Opens and closes every session — sets the focus, captures what was learned, enforces the governance rules, and hands off cleanly to the next one.
Intake
The filing clerk: triages anything received, proposes a destination, flags sensitivity, detects duplicates — and moves nothing without approval.
Job
Runs a piece of incoming work through the system without it becoming part of the system: an isolated, local-only workspace with the deliverable and the finish line agreed up front, closed to a one-line record.
Housekeeping
The acting counterpart to the read-only health check: finds loose ends — an unclosed session, a stranded commit, a silent agent — and fixes the pre-authorized classes, surfacing the rest.
Status
Read-only 'where am I' snapshot — composes the health checks and the governance files into one traffic-light view of what's active and what's next.
Mac health
The same read-only discipline applied to the machine itself — battery, storage, memory, what auto-starts. Reports; never changes a setting.
Tidy
File hygiene for the workspace — dead links, drifted metadata, stale records. Proposes every fix; moves nothing without approval.
Simplify
Re-renders a deep technical stretch of conversation back into plain English.
Diagnostic pipeline
Taking a raw idea to a verdict — the agents that do the thinking.
Diagnostic lab
Structured early-stage ideation and diagnosis; every idea that passes through is indexed for later reference.
Diagnostic agent
An autonomous orchestrator that runs the full cycle — live research, debate, specialist advisors, strategy — into a single compiled verdict.
Brainstorm team
Five distinct voices debate an idea through structured rounds — built for genuine tension, not polite agreement.
Specialist advisors
The panel the pipeline convenes — each pressure-tests one dimension before anything is built.
Commercial strategist
Positioning and go-to-market in one voice; diagnosis before tactics.
Financial analyst
Unit economics, break-even, pricing and runway — money in numbers, not vibes.
Technical advisor
Buildable? And the smallest thing that tests the riskiest assumption.
Distribution realist
Is there a credible path to customers? Names the distribution problem before anyone starts building.
Legal & regulatory advisor
A structured risk scan (UK-focused) that flags what needs professional attention before time or money is committed.
UX / customer-experience advisor
Will the first use actually reach value before the customer drops off?
Deep researcher
Grounds an idea in market reality — size, demand, competition — and gives a blunt call.
Knowledge & retrieval
The RAG and learning layer — grounded, cited answers over my own corpus.
RAG
Question-answering over the operator-workspace corpus, with grounded, cited answers.
Learn-ask
The same, over the separate learning corpus.
Learn-extract
Turns a podcast, talk, article or paper into a structured episode record.
Learn-bundle
Synthesizes a set of episodes into a single cross-episode bundle.
Learn-course
Captures a course as a structured record, mapped to the capability matrix.
Deepreach
Multi-model deep research: the same question fanned out across different models, verified, then consolidated — cross-model divergence is treated as a signal to check primary sources, not noise.
Quality gates
Nothing ships unreviewed.
Communication review
Scores a written artefact against quality standards and fixes the AI-tells.
Verify claims
Checks a document's factual claims against real sources and returns a verdict ledger — verified, unsupported, or incorrect — with the material claims checked hardest.
Roster verify
Refreshes a whole list of people or companies against live sources: reconciles the list against itself first, then returns each entry with what changed, how solid the evidence is, and where it came from. An optional second pass tries to disprove its own conclusions.
Site health
A build, security, pages and data-quality check across projects.
Measurement & profile
Keeping the capability picture honest.
Skills scan
Scans the workspace for the evidence behind each capability claim — and flags where I'm over- or under-claiming.
Skills export
Audience-shaped career outputs (CV, LinkedIn, bio, capability statement) generated from one source of truth.
Quiz
A one-question-per-session learning drip over how the workspace itself works — answers feed the capability matrix as corroborating evidence.
Built and refined in daily use — the toolkit grows as the work does.