CEO Agent — A Multi-Agent Outreach System Built on Claude
Role: Solo builder — architecture, agent design, prompt engineering, deployment Domain: AI automation for B2B sales development Stack: Claude (Anthropic API) · OpenClaw agent runtime · Python · JSON state layer · Linux VPS Repository:https://github.com/hajraby24-hue/ceo-agent

Summary
CEO Agent is an autonomous business-development system that runs an entire outbound sales motion — from finding prospects to sending personalised messages to tracking who replied — without a human sitting in the loop for each step.
Rather than one large prompt trying to do everything, the system is built as an orchestrator with nine specialised sub-agents. A "CEO" layer decides what needs to happen next; each sub-agent owns one narrow job and hands its output to the next. The result is a system that behaves less like a chatbot and more like a small sales team with a defined org chart.
This case study focuses on how it was designed and built rather than on campaign results.
The Problem
Outbound prospecting for small and mid-sized businesses is repetitive but not simple. A single campaign requires you to:
Find businesses that fit a profile in a specific market
Verify they actually exist, are active, and are reachable
Decide whether each one is worth contacting
Write a message that doesn't read like a template
Send it on the right channel at the right time
Track state so nobody gets messaged twice
Know what to do when someone replies
Most automation tools handle step 5 well and everything else badly. Generic LLM wrappers handle steps 4 and 7 well and everything else badly. The gap is orchestration — knowing which step you're on, for which prospect, and what the last step produced.
CEO Agent was built to close that gap.
Architecture
Design principle: the workspace is the agent
The system doesn't live in a single application file. It lives in a workspace directory where three kinds of files sit side by side:
Layer | Format | Purpose |
Identity & behaviour | Markdown | Who the agent is, how it reasons, what it must never do |
Execution | Python | Deterministic actions — search, send, log, retry |
State & memory | JSON | What has happened, to whom, and when |
This separation matters. Reasoning belongs in natural language, actions belong in code, and truth belongs in structured state. Mixing them is the most common reason agent systems become unmaintainable.
Core definition files include:
IDENTITY.md — the agent's role, mandate, and scope of authority
SOUL.md — tone, judgment principles, and behavioural constraints
AGENTS.md — the roster and the handoff contract between agents
MEMORY.md — persistent context that survives across sessions
Because these are plain Markdown, the system's behaviour can be changed by editing prose — no redeployment, no code changes, and the diff is readable by a non-engineer.
The agent roster
The CEO Agent orchestrates nine specialists. Each has a single responsibility and a defined input/output contract:
Agent | Responsibility |
Research | Discovers candidate businesses and enriches them with public data |
Leads | Normalises raw findings into structured lead records |
Qualification | Scores fit and filters out prospects not worth contacting |
Outreach | Drafts channel-appropriate messages personalised to each lead |
Sales | Handles conversation logic once a prospect responds |
Pipeline | Tracks stage, ownership, and next action per lead |
Recommendation | Suggests the highest-value next move across the whole pipeline |
QA | Reviews drafted output before anything leaves the system |
Output | Formats and dispatches the final artefact to the channel |
The QA agent is the design decision I'd defend hardest. Autonomous outreach fails publicly — a broken merge field or a malformed number is visible to a real customer. Inserting a review step between drafting and sending trades a small amount of latency for a large reduction in embarrassing failures.
Orchestration flow
CEO Agent
│
├─▶ Research ──▶ Leads ──▶ Qualification
│ │
│ ▼
│ Outreach ──▶ QA ──▶ Output ──▶ [WhatsApp / Email]
│ │
│ ▼
└────────────── Pipeline ◀── Sales ◀────────────────────── Response
│
▼
Recommendation ──▶ back to CEO AgentThe loop is deliberate. Pipeline state feeds the Recommendation agent, which feeds the orchestrator, which decides what the next cycle should prioritise. The system gets more useful the longer it runs because its state file is richer.
Implementation Details
Model selection
The system runs on Claude Haiku for the high-frequency, well-scoped tasks — classification, normalisation, template filling, QA checks. These calls happen hundreds of times per campaign and don't require deep reasoning, so latency and cost per call dominate the decision.
The architectural point: model choice is a per-agent decision, not a system-wide one. Because each sub-agent is a separate call with a separate prompt, a heavier model can be swapped into the Outreach or Sales agent — where message quality genuinely moves the outcome — without changing anything else.
State management
Every send is written to an append-only JSON log before the next action is attempted. This makes the system:
Idempotent — a campaign can be re-run without double-messaging anyone
Resumable — a crashed run picks up where it stopped
Auditable — every message sent can be traced to a lead, a timestamp, and a template version
Separate logs track outreach attempts, WhatsApp delivery, and overall campaign state, so a failure in one channel doesn't corrupt the record of another.
Batching and retry
Campaign execution is split into batch scripts rather than one long-running process. A resend script handles the tail of failures separately from the main run. This is unglamorous but it's what makes the difference between a demo and something you can leave running.
Message templating
Outreach copy lives in versioned Markdown template files, localised for the target market rather than translated into it. Template selection is an agent decision; template content is human-authored and reviewed. Fully generative outbound copy was tested and rejected — the variance was too high for messages that represent a business.
Deployment
The system runs on a Linux VPS under the OpenClaw agent runtime, which provides the process supervision, credential management, and channel adapters. The entire workspace — identity files, scripts, state, and logs — is self-contained and portable, which means the full agent can be archived, versioned, and redeployed as a single directory.
That portability was a deliberate constraint from day one: an agent you can't move is an agent you don't own.
Engineering Challenges
Keeping agents in their lane. Early versions had sub-agents quietly re-doing each other's work — the Outreach agent would re-qualify a lead it had been handed. Fixed by tightening the handoff contract in AGENTS.md and making each agent's input schema explicit rather than conversational.
State drift between channels. A lead contacted by email and separately by WhatsApp appeared as two records. Fixed by making the lead identity key channel-independent and reconciling logs against a single pipeline record.
Failure visibility. Autonomous systems fail silently by default. Adding structured logging at every dispatch point meant failures showed up as data rather than as an absence of results.
Responsible Use
Outbound automation touches real people's contact details, so the system was built with hard constraints:
Contact data is sourced from public business listings, not scraped personal profiles
Suppression state is checked before every send; opt-outs are permanent
Send volume is rate-limited by design, not by accident
Campaign logs containing contact information are treated as sensitive and stored accordingly
Automating outreach is not the same as automating away judgment about who should be contacted and how often.
What This Project Demonstrates
Multi-agent system design — decomposition, handoff contracts, and orchestration loops
Production LLM engineering — model selection per task, cost-aware architecture, QA gating
Prompt architecture as configuration — behaviour defined in editable natural language, not buried in code
Operational discipline — idempotency, resumability, audit trails, and retry handling
Domain integration — applying agent architecture to a real go-to-market problem rather than a toy task
Roadmap
Per-agent model routing with a heavier model on the Outreach and Sales layers
A reporting layer surfacing pipeline state as a live dashboard
CRM write-back so the agent's pipeline syncs with a system of record
Evaluation harness scoring Outreach drafts against a human-rated reference set
Built by Mohammed Yasir — Digital Marketing Specialist & AI Automation Engineer www.mohammed-yasir.com



Comments