top of page
Search

CEO Agent — A Multi-Agent Outreach System Built on Claude

Writer: Mohammed Yasir
Mohammed Yasir
Aug 29
5 min read

Role: Solo builder — architecture, agent design, prompt engineering, deployment Domain: AI automation for B2B sales development Stack: Claude (Anthropic API) · OpenClaw agent runtime · Python · JSON state layer · Linux VPS Repository:https://github.com/hajraby24-hue/ceo-agent


Summary

CEO Agent is an autonomous business-development system that runs an entire outbound sales motion — from finding prospects to sending personalised messages to tracking who replied — without a human sitting in the loop for each step.

Rather than one large prompt trying to do everything, the system is built as an orchestrator with nine specialised sub-agents. A "CEO" layer decides what needs to happen next; each sub-agent owns one narrow job and hands its output to the next. The result is a system that behaves less like a chatbot and more like a small sales team with a defined org chart.

This case study focuses on how it was designed and built rather than on campaign results.


The Problem

Outbound prospecting for small and mid-sized businesses is repetitive but not simple. A single campaign requires you to:

  1. Find businesses that fit a profile in a specific market

  2. Verify they actually exist, are active, and are reachable

  3. Decide whether each one is worth contacting

  4. Write a message that doesn't read like a template

  5. Send it on the right channel at the right time

  6. Track state so nobody gets messaged twice

  7. Know what to do when someone replies

    Most automation tools handle step 5 well and everything else badly. Generic LLM wrappers handle steps 4 and 7 well and everything else badly. The gap is orchestration — knowing which step you're on, for which prospect, and what the last step produced.

    CEO Agent was built to close that gap.


Architecture

Design principle: the workspace is the agent

The system doesn't live in a single application file. It lives in a workspace directory where three kinds of files sit side by side:

Layer

Format

Purpose

Identity & behaviour

Markdown

Who the agent is, how it reasons, what it must never do

Execution

Python

Deterministic actions — search, send, log, retry

State & memory

JSON

What has happened, to whom, and when

This separation matters. Reasoning belongs in natural language, actions belong in code, and truth belongs in structured state. Mixing them is the most common reason agent systems become unmaintainable.

Core definition files include:

  • IDENTITY.md — the agent's role, mandate, and scope of authority

  • SOUL.md — tone, judgment principles, and behavioural constraints

  • AGENTS.md — the roster and the handoff contract between agents

  • MEMORY.md — persistent context that survives across sessions

Because these are plain Markdown, the system's behaviour can be changed by editing prose — no redeployment, no code changes, and the diff is readable by a non-engineer.


The agent roster

The CEO Agent orchestrates nine specialists. Each has a single responsibility and a defined input/output contract:

Agent

Responsibility

Research

Discovers candidate businesses and enriches them with public data

Leads

Normalises raw findings into structured lead records

Qualification

Scores fit and filters out prospects not worth contacting

Outreach

Drafts channel-appropriate messages personalised to each lead

Sales

Handles conversation logic once a prospect responds

Pipeline

Tracks stage, ownership, and next action per lead

Recommendation

Suggests the highest-value next move across the whole pipeline

QA

Reviews drafted output before anything leaves the system

Output

Formats and dispatches the final artefact to the channel

The QA agent is the design decision I'd defend hardest. Autonomous outreach fails publicly — a broken merge field or a malformed number is visible to a real customer. Inserting a review step between drafting and sending trades a small amount of latency for a large reduction in embarrassing failures.


Orchestration flow

CEO Agent
    │
    ├─▶ Research ──▶ Leads ──▶ Qualification
    │                              │
    │                              ▼
    │             Outreach ──▶ QA ──▶ Output ──▶ [WhatsApp / Email]
    │                                                             │
    │                                                             ▼
    └────────────── Pipeline ◀── Sales ◀────────────────────── Response
                       │
                       ▼
                 Recommendation ──▶ back to CEO Agent

The loop is deliberate. Pipeline state feeds the Recommendation agent, which feeds the orchestrator, which decides what the next cycle should prioritise. The system gets more useful the longer it runs because its state file is richer.


Implementation Details

Model selection

The system runs on Claude Haiku for the high-frequency, well-scoped tasks — classification, normalisation, template filling, QA checks. These calls happen hundreds of times per campaign and don't require deep reasoning, so latency and cost per call dominate the decision.

The architectural point: model choice is a per-agent decision, not a system-wide one. Because each sub-agent is a separate call with a separate prompt, a heavier model can be swapped into the Outreach or Sales agent — where message quality genuinely moves the outcome — without changing anything else.

State management

Every send is written to an append-only JSON log before the next action is attempted. This makes the system:

  • Idempotent — a campaign can be re-run without double-messaging anyone

  • Resumable — a crashed run picks up where it stopped

  • Auditable — every message sent can be traced to a lead, a timestamp, and a template version

Separate logs track outreach attempts, WhatsApp delivery, and overall campaign state, so a failure in one channel doesn't corrupt the record of another.

Batching and retry

Campaign execution is split into batch scripts rather than one long-running process. A resend script handles the tail of failures separately from the main run. This is unglamorous but it's what makes the difference between a demo and something you can leave running.

Message templating

Outreach copy lives in versioned Markdown template files, localised for the target market rather than translated into it. Template selection is an agent decision; template content is human-authored and reviewed. Fully generative outbound copy was tested and rejected — the variance was too high for messages that represent a business.


Deployment

The system runs on a Linux VPS under the OpenClaw agent runtime, which provides the process supervision, credential management, and channel adapters. The entire workspace — identity files, scripts, state, and logs — is self-contained and portable, which means the full agent can be archived, versioned, and redeployed as a single directory.

That portability was a deliberate constraint from day one: an agent you can't move is an agent you don't own.


Engineering Challenges

Keeping agents in their lane. Early versions had sub-agents quietly re-doing each other's work — the Outreach agent would re-qualify a lead it had been handed. Fixed by tightening the handoff contract in AGENTS.md and making each agent's input schema explicit rather than conversational.

State drift between channels. A lead contacted by email and separately by WhatsApp appeared as two records. Fixed by making the lead identity key channel-independent and reconciling logs against a single pipeline record.

Failure visibility. Autonomous systems fail silently by default. Adding structured logging at every dispatch point meant failures showed up as data rather than as an absence of results.


Responsible Use

Outbound automation touches real people's contact details, so the system was built with hard constraints:

  • Contact data is sourced from public business listings, not scraped personal profiles

  • Suppression state is checked before every send; opt-outs are permanent

  • Send volume is rate-limited by design, not by accident

  • Campaign logs containing contact information are treated as sensitive and stored accordingly

Automating outreach is not the same as automating away judgment about who should be contacted and how often.


What This Project Demonstrates

  • Multi-agent system design — decomposition, handoff contracts, and orchestration loops

  • Production LLM engineering — model selection per task, cost-aware architecture, QA gating

  • Prompt architecture as configuration — behaviour defined in editable natural language, not buried in code

  • Operational discipline — idempotency, resumability, audit trails, and retry handling

  • Domain integration — applying agent architecture to a real go-to-market problem rather than a toy task

Roadmap

  • Per-agent model routing with a heavier model on the Outreach and Sales layers

  • A reporting layer surfacing pipeline state as a live dashboard

  • CRM write-back so the agent's pipeline syncs with a system of record

  • Evaluation harness scoring Outreach drafts against a human-rated reference set


Built by Mohammed Yasir — Digital Marketing Specialist & AI Automation Engineer www.mohammed-yasir.com

 
 
 

Comments


©2026 by Mohammed Yasir. Powered and secured by Wix

bottom of page