Deadwater

sept 18 2026 · updated sept 19 2026

By Jack Virag

Pick your first two AI workflows before you buy the platform

Compare real marketing jobs by input readiness, review effort, failure cost, and ownership. Use a small pilot to choose the next workflow and the right tooling.

9 min read
ai-workflowsmarketing-operationscontext-osautomation
Pick your first two AI workflows before you buy the platform

“Automate marketing” is a fantastic way to buy software and inherit a second job.

The phrase hides research, decisions, production, review, and publishing inside one attractive promise. A platform demo can make that whole bundle look like a button.

Then your team has to operate the button.

Before comparing platforms, pick a first workflow and a provisional second. Pilot the first. Let what you learn confirm or replace the second.

Two gives you enough room to think beyond a party trick without pretending you're redesigning the whole department. The number is a scoping choice. The important part is choosing the work before buying its container.

Find the job under the ambition

“Research our customers” could mean almost anything. “Turn one approved interview transcript into a source-linked brief for the product marketer” gives you something a person can accept or reject.

Start there. Find work the team already does, with a recognizable input and someone waiting for the result.

Google's People + AI Guidebook recommends mapping existing work and deciding whether AI adds value. Watch the work before imagining its replacement. Collect a few recent examples, including the annoying one.

Follow the last actual instance

Ask the operator what triggered the job, what they gathered, where they got stuck, and what they handed over. Keep hands-on time separate from waiting.

A brief that arrives two days late may require 30 minutes of writing and a day and a half in a product-approval queue. Automating the writing leaves most of that problem intact.

Frequency matters, but so does use. A daily report nobody reads is a poor candidate for a more elaborate production system.

Google's SRE guidance on operational toil makes another useful distinction: operating a script can still take manual work. And work requiring essential human judgment isn't automatically waste.

Copying the same fields between tools may be toil. Deciding whether a customer story is fair and compelling may be why you hired an editor.

You can run a fuller content-system audit, but choosing a pilot doesn't require inventorying the entire company. Get enough evidence to describe a few real jobs.

Put one job on a card

For the interview example, the card could be this simple:

Question Illustrative answer
What starts it? An approved transcript enters the queue
What does it need? Transcript, current product facts, brief template, good example
What comes out? One source-linked internal brief with open questions
Who judges it? The designated product marketer
What must it preserve? Meaning, source support, current product names
What can't it do? Publish or contact the interview participant

Add what happens when an input is missing. A transcript without speaker labels may need correction. An unsupported product claim should remain a question.

A useful brief carries the decision the next person needs to make. It isn't just a topic with a pile of instructions stapled to it.

If that small output removes the bottleneck, stop the first workflow there. Drafting, distribution, and reporting can wait, even if the platform has very appealing buttons for all three.

Don't make the model do the plumbing

Moving an approved file, checking a required field, and applying a naming convention may need ordinary code. Interpreting an interview may benefit from a model.

Anthropic's agent architecture guidance distinguishes predefined workflows from systems that choose their own path, and recommends starting simply.

That distinction is more useful than asking whether the platform is sufficiently “agentic.” What decisions does this particular job need to make at runtime? Which steps can you already write down?

Keep the known steps boring. Save the model's judgment for the part that needs it.

Score the work you can start, not the future you can pitch

“This could save our entire team” is a hypothesis that needs much better handwriting.

For a first pilot, readiness deserves more attention than imaginary upside. You need usable inputs, someone to operate the workflow, and a way to judge whether its output helped.

First apply a few gates:

  • You have permission to use the source material.
  • An operator has time and authority to run the pilot.
  • A reviewer can describe an acceptable result.
  • The workflow's permitted actions are clear.

A high score elsewhere cannot compensate for a failed gate. If nobody can review the output, “human in the loop” is a staffing vacancy with a diagram.

Make the disagreements visible

The following scorecard is a discussion tool, not a calibrated prediction of savings. Give each dimension zero, one, or two based on evidence from the actual job.

Dimension Zero One Two
Frequency Rare or unclear demand Recurring, limited examples Enough repetition to observe use
Inputs Missing or unapproved Available, need cleanup Permitted, current, usable
Review Unknown acceptance bar Defined, substantial review Bounded review with a clear rubric
Failure cost Consequential action without workable controls Recoverable with cleanup Contained draft with inspectable corrections
Owner capacity Nobody accountable Named, capacity unresolved Named, with time and authority

Have the operator and reviewer score independently. If the builder gives review a two and the editor gives it a zero, open an example together.

The disagreement is the useful part. It may expose missing context, a bad output definition, or a reviewer who expects to reconstruct every draft.

AirOps has a documented Human Review step that can pause work for editing, acceptance, or cancellation. A platform can provide that state. It cannot provide the person's calendar.

Give the second workflow a pencil booking

Imagine a team with usable interview transcripts, current product facts, an editor, and an article backlog whose sources need cleanup. These invented conditions might produce this comparison:

Candidate Frequency Inputs Review Failure Owner Decision
Transcript to internal brief 1 2 2 2 2 9: first pilot
Article-refresh preparation 2 1 1 2 2 8: provisional second
Unreviewed article publication 2 1 0 0 1 4: outside this pilot's permitted scope

The transcript job wins because its inputs and review task are already in decent shape. Refresh preparation has more available work, but more source ambiguity too.

That second workflow could produce a proposed change list with evidence and unresolved questions. It doesn't have to rewrite and publish the entire backlog to be useful.

Microsoft's guidance to scope services when uncertain supports clarification or fallback when the intended action is unclear. For a refresh workflow, “this product claim needs an owner” can be the right output.

Map those decisions into the actual workflow stages. Source checking, editorial acceptance, and publication permission answer different questions. If publishing enters the scope later, give its QA gate its own conditions.

From a good idea to a working system

A Context OS connects your company knowledge to repeatable work. See what goes into one.

The pilot has to be allowed to disappoint you

If the second workflow is locked into the package before the first runs, the pilot is mostly theater for the purchasing process.

Set a duration, input set, acceptance bar, and decision date. Include the awkward cases you already know about. Hold some examples back during development so the test goes beyond familiar demonstrations.

Google's data and evaluation guidance emphasizes real-world conditions and unseen data. Anthropic's agent-evaluation guidance also distinguishes claimed success from the actual result.

Open the saved brief. Check its sources. Have the receiving person judge whether they can use it. Repeat important cases instead of keeping only the best attempt.

Record enough to compare the pilot with the current process:

  • Preparation time.
  • Review and rework time.
  • Elapsed time, including waiting.
  • Service spend and remaining manual steps.
  • Accepted, rejected, and pending outputs.
  • Critical failures and the input that exposed them.

Keep setup effort separate from recurring work. Use similar input difficulty and the same acceptance bar on both sides of the comparison.

GOV.UK's guidance on measuring service benefits starts with a baseline and realistic estimates, including less favorable outcomes. Rejected drafts and extra checking belong in your version of that record.

Faster delivery can be useful without saving staff time. Better coverage can be useful without reducing cost. Name the benefit you observed instead of converting every good thing into a fictional savings figure.

Then let the result change the plan.

If the briefs are useful and review fits the team's capacity, refresh preparation may make a sensible second workflow. It can reuse some sources and conventions while adding its own tests.

If the editor spends most of the time hunting missing context, another drafting workflow feeds the same bottleneck. The second job might prepare evidence for review. Or the next project might be getting product facts into a maintainable state.

For a simple capacity example, eight briefs at 20 minutes of review each create 160 minutes of work. An editor with 90 available minutes is already over capacity. A faster generation step doesn't close that gap.

Make the platform audition for the job

Now you can bring something useful to the demo: representative inputs, an accepted output, a missing fact, a rejected result, and a retry.

Ask the vendor to walk through those cases. Inspect where the work lands, what the operator can change, and what requires additional implementation.

Our workflow software buyer guide covers the broader landscape. Your job cards make that comparison specific. Sometimes the right first implementation belongs in software you already have.

When you scope the first two workflows with Deadwater, bring recent examples and the person who reviews the work. We can start with a job they actually want done.

The platform should earn its place in that plan. It shouldn't get to write the plan because its demo was good.

Put this to work

Bring a recurring content problem. We’ll help scope the system behind it.