I build the agent teams behind this work, and the loop that improves them
Agent teams governed by a single rubric, bounded loops with hard fact-check gates, and a meta-optimizer whose judges stay frozen and may only get stricter. One skeleton, archived on the US site build and the CMS rebuild, in progress on three more.
One pattern, six marketing functions
Agent team + loop + rubric + a human on the gate, pointed at one marketing function until it runs itself.
The US site build used it as a build pipeline. The CMS rebuild was filled by this content team. The reporting model is meant to fold its Layer 4 onto this engine, and that one has not happened yet. The same skeleton is what I am extending to Google Ads ops, SDR outreach and marketing email; none of those has an archived run yet.
This page is the machinery; the work it produced is credited on those project pages.
I am rebuilding a team I used to coordinate
At PingPong the website was a human operation. Webflow, and around it a product marketing manager, a product manager, a content manager, a designer and a content writer, all working the same page. Then the review: the marketing team ran several sprints where every person reviewed at least their own page, wrote up their feedback, and we discussed what to implement before anything changed. It was systematic and it worked. It was also slow, and it consumed most of a team.
At ThinkingAI I own website operations, inbound campaigns and the tech stack at the same time, so that shape was never going to fit. So I took the collaboration apart: every seat in that old room became an agent with its own context, its own skills, and its own to-do list, and the sprint review became a review loop. The systems below are that team taken apart, and a new page now comes out in a few hours.
The agent roster is a copy of a real org chart: product marketing, writer, designer, UI/UX, assembly, reviewers, plus a loop that improves themFour systems: a page factory, a review loop, a meta-optimizer, a content team
All four exist in a real repository with archived runs. One builds pages from a module catalog, and one reviews the copy. A third writes and ships content with its own images, and the fourth improves the other three. The same skeleton runs behind the rest of the Resource Center, where every client-facing content type has its own team that produces a piece, assembles it, reviews it, illustrates it and publishes it, with a final visual check.
A catalog of 30+ modules, and agents that pick from it
I started by taking inventory. Each of the more than 30 reusable module templates I built declares its own constraints: which scenarios it suits, where it belongs on a page, roughly how much text each section holds, and the image dimensions it expects. That turns layout into a catalog an agent can choose from instead of a blank canvas it has to invent against.
A page run then works the way the old team did. The product marketing agent carries the product-capability knowledge base as its context. The writer agent carries brand tone of voice and the writing skills, held as system prompts. A UI/UX agent decides which module each block of content needs and picks the scheme from the catalog. The designer is its own sub-team: a visual instruction goes out, an image tool generates the asset, Claude Code inspects what came back, and an agent composites the HTML overlay when the module calls for one. An assembly agent puts the page together at the end.
Every module template predeclares scenario, placement, text budget per section and image dimensions, so the agent picks from a listDraft in, launch-ready out
An 8-stage pipeline takes raw page copy to publishable. A context distiller builds a cacheable brief, two prospect personas and a CEO persona score the draft in parallel, a synthesis stage resolves their conflicts into one change brief, and an architect revises. Then the hard part: a fact-check gate where any claim without a source is a must-fix, a de-AI pass, and a re-score by the same three reviewers against baseline. A human holds the publish gate.
The three reviewers run on three different vendors' frontier models, each briefed as a first-time visitor from our ICP with no prior context, asked the same blunt question: how fast can you tell what this is. They score, they read each other's findings, and they argue. Three vendors means one vendor's blind spot stays one vendor's blind spot.
8 personas defined in markdown · one scoring rubric as the single source of truth · every run archived in fullThe loop that proposes fixes to the team's own rules
This loop's job is making the team better. It traces every deduction to one root cause, then splits them: I can fix the instructions, the context and the rubric. Missing inputs, human decisions and real tradeoffs stay where they are. It proposes diffs only for the fixable kind, and an adversarial guardrail interrogates each one: is it a real defect or score-gaming, does it generalize, and what are the side effects. Survivors come to me for review. Nothing auto-applies.
A quieter kind of improvement runs alongside it. Each page run has checkpoints where an agent can write back what it learned, so the skills and knowledge an individual agent needed on this page persist into its long-term memory for the next one. The meta-optimizer changes the team's rules through a human gate. This just means each agent starts the next page knowing more than it did on the last.
The score tells me where to look; I never let the system optimize against it. Judges stay frozen and may only get stricter, and adopted changes are re-validated on a different pageThe one diff the guardrail sent back
In the 2026-07-06 run, all three reviewers flagged the same phrase: “rented LLM,” jargon sitting on the site's strongest differentiator. The obvious fix, diff D8, would let the final Content Master pass rewrite awkward terms into plain English. The guardrail caught what a naive optimizer would ship: the Content Master runs at stage 50, and the fact-check gate runs at stage 40, on the earlier draft, so a blanket rewrite permission would let fact-constrained claims land in final copy with no fact-check behind them. Verdict: REVISE. Ordinary jargon may be rewritten freely; any term touching an honesty red line may only be proposed by the Content Master.
Across all 10 diffs that round: every proposal cited a real failure, none loosened a gate · 9 ADOPT · 1 REVISE · 0 REJECTRaw material to published asset, images included
Five agents take source material to a live CMS entry. Product Marketing writes the guideline, outline, and image spec. The Writer drafts under hard de-AI rules, and a Reviewer scores each piece through a dual-lens gate and leaves a record. Then the Designer generates the 16:9 cover via Volcano Ark Seedream 4.0, and Assembly builds the Sanity document and publishes it through the HTTP mutation API.
The same skeleton already runs product updates, blog covers, and a version tuned for case studiesWhere this actually stands
Everything exists as archived runs and human-gated proposals in a real repository. This page has no external track record yet.
No external numbers yet
The quality and speed lift is internal; I have no hard external outcome metrics yet, and I mark them to-be-measured rather than imply them. The content these systems produced is credited to the CMS Rebuild, with no double counting. This page claims the machinery, and only the machinery.
Reusability
Personas live in markdown, one rubric scores everything, and a separate script runs the loop. The system retunes to a new page or a new content type by editing markdown, with no rebuild.
Related work
Claude Code
Seedream 4.0
Sanity API
Frozen judges, bounded loops, a human on publish
Hill-climbing the judges' score converges on copy that flatters the judges. Goodhart's law, in production. So the score stays an instrument panel, and I never maximize it.
The review personas never get easier to please. Any change to a judge must tighten it. I only change the side that writes, which keeps the delta between runs meaningful.
No diff gets in on plausibility. Each diff must point at a documented failure in the run artifacts. In the archived round, all ten did.
What agents say lives in markdown; how the loop runs lives in one workflow script; every score traces to one rubric file. Retuning it is a markdown edit. So far I am the only person who has done that.
Revision loops have a fixed budget, so they cannot churn forever. The fact-check gate is a hard stop: a claim without a source is a must-fix, with one bounded patch. Outputs are schema-validated.
The machine does the legwork end to end, then stops. Publishing a page and adopting a system change both require a person to say yes. The one ungated path is agent memory.
Want this machinery pointed at your function?
Its rules are markdown, so pointing it at a new function is an editing job.