Deadwater

mar 30 2026 · updated sept 19 2026

By Jack Virag

Stop fixing the same AI draft twice

Turn repeated AirOps editorial corrections into scoped rules, better examples, and source fixes that you can test on the next draft.

10 min read
airopseditorial-workflowscontent-opscontext-os
Stop fixing the same AI draft twice

The editor fixed that claim yesterday. Today, the AI has thoughtfully put it back.

You can spend a surprising amount of time “improving” a content system by making the same correction in every draft it produces.

The article gets better. The system learns absolutely nothing from the transaction because nothing it uses next time changed.

In AirOps, the missing connection is buildable. Capture the reason for the edit, find where the error came from, change the relevant maintained input, and test another draft. The trick is making that correction reusable without turning every preference into a permanent law.

Save the reason, not just the replacement sentence

Consider a fictional support product called Queueworks. It suggests replies that a support representative reviews and sends.

The draft says, “Queueworks automatically closes every ticket.” The editor replaces that with the accurate capability.

If you save only the replacement, a later process might conclude that the editor dislikes “automatically.” Or that the preferred sentence length is longer. Neither captures what went wrong.

The draft turned an assistive capability into a completed outcome. That's the reason to preserve alongside the original passage, correction, asset, revision, and product source.

AirOps' Human Review step pauses execution for editing or selection, then returns values after acceptance or cancels the run. Later steps can use those values.

Accepting an edit doesn't automatically install a persistent rule or train the model for future work. If you need the reason, collect it deliberately.

An editor can write a short explanation. A model can propose a classification for confirmation. When the reason is unclear, preserve the question rather than confidently routing the wrong diagnosis.

The full AirOps writing workflow needs that information at the revision boundary. “Make this better” is an extremely portable way to lose an editor's judgment.

Inspect the input before writing another rule

Was the approved product source loaded? Was it current? Did an old sample contradict it? Did the writing step receive the intended product line at all?

Suppose the accurate specification was present, but an old sample guide contained the ticket-closing claim. That sample is a concrete defect to repair. It's a plausible contributor, not proof of the sole cause.

Keep the failed input so you can test the diagnosis.

Finding Likely place to fix it
Product record is wrong Maintained factual source, with its owner
Example teaches obsolete behavior The example
Correct source never reached the writer Retrieval or input configuration
Recurring format-specific mistake Content-type guidance and examples
One article's opening needs a different angle That article

A known factual error doesn't need to happen three times before it deserves repair. Conversely, a one-off preference doesn't become a universal rule just because it arrived in a comment.

AirOps Content Types can hold format-specific rules and sample URLs. For Queueworks implementation guides, a useful instruction might distinguish the software's action from the user's action.

Keep the actual capability in the product source. The example should demonstrate the distinction, not become a second product specification.

Otherwise you're paying the prompt brittleness tax by appending another global warning every time an editor touches a sentence.

Give the current edit a precise route

The paragraph still needs fixing. A structured handoff helps the revision step know what to change, what to retrieve, and what to preserve.

This fictional example conforms to the downloadable editorial feedback schema. It's a proposed integration object, not an AirOps payload that appears automatically after review.

{
  "status": "needs_revision",
  "review_id": "review-example-08",
  "asset_id": "queueworks-implementation-guide",
  "revision_scope": "paragraph",
  "items": [
    {
      "id": "feedback-01",
      "type": "factual_accuracy",
      "section": "How suggested replies work",
      "instruction": "Describe suggested replies and the support representative's review-and-send role using the approved product specification.",
      "severity": "high",
      "status": "open",
      "retrieve_from": ["product_truth"],
      "preserve": ["Unrelated approved sections", "Existing supported examples"],
      "acceptance_checks": [
        "The capability matches the current approved specification",
        "The representative's role remains explicit",
        "No unsupported ticket-closing claim remains"
      ]
    }
  ]
}

Validate the object before routing it. JSON Schema's object rules support required fields and restrictions on extra properties.

The download requires asset identity, scope, and feedback items, with allowed categories and retrieval values. Those checks can reject a malformed handoff. They can't decide whether Queueworks closes tickets.

The schema also distinguishes addressed from verified. Give those states meaning in the workflow: a replacement can exist before anyone has checked its factual accuracy.

The schema alone doesn't enforce that transition or prevent an approved root object from containing unresolved items. Your routing and release conditions need to do that work.

Different feedback deserves different revision paths

The map below illustrates the routing idea. Click a category to see a predefined note, route, and acceptance checks. It doesn't capture feedback, save changes, validate JSON, or connect to AirOps.

Interactive revision map

Click the feedback type and watch the workflow route change

Editorial comments should not all trigger the same rewrite path. Different kinds of feedback need different retrieval, revision, and QA behavior.

Download JSON schema

Reviewer note

This product claim is too broad. Re-ground it in source truth.

Workflow route

  1. 1Read product-truth source
  2. 2Revise the affected section only
  3. 3Preserve approved sections
  4. 4Run factual QA before approval

Checks before approve

  • Source truth loaded
  • Claim narrowed
  • No new unsupported language

To build it, connect review capture, object validation, relevant source retrieval, targeted revision, and result checking. Store the outcome against the article.

For the Queueworks correction, load the product specification and revise the affected paragraph. Preserve unrelated approved work. If another paragraph repeats the same false promise, expand the scope deliberately to fix that too.

AirOps' Content Comparison step can expose the changes. Its inputs require HTML, so convert Markdown first. The documented enhanced Grid accept/reject view depends on placing the comparison as the final workflow step.

A visible diff helps the reviewer inspect the edit. It doesn't supply the factual judgment.

Give the pre-publish QA gate specific failure paths. Missing identity should stop processing. Missing evidence should return a source question. An awkward transition should get an editorial edit.

Sending all three through “rewrite this article” throws away the useful part of the diagnosis.

From a good idea to a working system

A Context OS connects your company knowledge to repeatable work. See what goes into one.

Make a second change for tomorrow's draft

Closing the paragraph's feedback item says nothing about the stale sample still waiting for the next run.

Track the maintained change separately. For our fictional example:

Field Entry
Scope Queueworks implementation guides describing suggested replies
Source decision Product owner confirms the existing specification
Example change Replace the ticket-closing sample
Instruction change Distinguish software action from user action
Owner Content owner maintains the example, guidance, and tests
Evidence Failed input, source revision, corrected sample, test outputs
State Candidate prepared; activation and subsequent-run check pending

Keep this outside the strict downloadable feedback object. Its contract doesn't include maintained-rule versions, owners, or test history. Adding those keys casually would make the object invalid.

AirOps Brand Kits have version history, comparison, and restoration. Draft edits remain non-live until published.

Record the candidate you reviewed and the revision you activated. An authorized owner can do that without convening another approval meeting for an already authorized correction.

Check what the writing step receives

AirOps uses a configured Brand Kit input and variable references, with selected product lines and content type, plus optional audience and region.

Inspect those selections and the resolved input. Did this run actually receive the corrected sample, applicable rule, and current facts?

The Brand Kit setup guide covers organizing the context. For this fix, the narrower question is whether the intended context reached the intended step.

A screenshot of the edited settings page doesn't answer that.

Test the candidate in an appropriate configuration before activation. If the test path can't select unpublished context, supply it explicitly in an isolated setup. Don't assume a live workflow reads a Brand Kit draft.

After the authorized activation, inspect a subsequent run. Include relevant neighboring workflows if they share the changed context, or narrow the change so they aren't accidentally affected.

The current article can be fixed while this durable change remains pending. Keep both states honest.

Test the fix against a case it shouldn't change

A rule that stops false automation claims by making every product sound manual is also a bad rule.

Anthropic's evaluation guidance recommends cases from observed failures, including when behavior should and shouldn't occur. For Queueworks, I'd start here:

Case Expected result
Original failed input Describe suggestions and the representative's review-and-send role
Fresh brief about the same capability Preserve the distinction without copying the saved paragraph
Documented automatic ticket tagging Keep the supported automation claim
Requested capability absent from the specification Flag the missing evidence
Feedback missing asset identity Reject before revision
Unrelated approved paragraph Preserve its meaning

Judge meaning, not exact wording. A banned-word check could catch “automatically” while accepting “handles the whole queue for you.” It could also reject the accurate tagging example.

AirOps supports step, branch, and whole-workflow tests. Check the current execution status as well as the visible values: its documentation warns that displayed step values from an all-step test don't update on error.

An old good output still sitting on screen is a particularly annoying way to congratulate yourself too early.

Count recurrence without hiding the rejects

Define which subsequent drafts qualify for review. In this example: first drafts of implementation guides describing suggested replies. Include eligible drafts that later get rejected.

The recurrence measure is affected eligible drafts divided by all eligible drafts reviewed. Six out of twelve is 50%; three out of twelve is 25%.

Those invented numbers explain the calculation. They aren't results, a recommended sample size, or evidence that one rule caused an improvement.

Keep topic mix, model, source and workflow versions, and review criteria beside the count. Replaying matched inputs helps, but add fresh cases so you're not tuning forever against yesterday's paragraph.

Track review and rework time too, including the effort spent maintaining the fix. Fewer comments don't help if the reviewer stopped checking or the workflow stopped saying anything useful.

The broader content QA process should catch new defects. If legitimate capabilities get understated, narrow the rule. Replacing the stale example may turn out to be enough.

Keep recovery specific

AirOps workflow versioning distinguishes restoring a historical version into the draft from setting a production default. Restore overwrites the current draft; Set Default leaves it intact.

Preserve unrelated work and identify the version you intend to run. That operation doesn't establish that the Brand Kit, external sources, or generated articles also reverted.

Inspect the compatible context and affected outputs separately.

Take one repeated correction through this loop before building a giant feedback library. You should be able to point to the reason, the maintained change, its owner, and what happened on the next relevant draft.

If that connection is missing, we can help build it. The editor should get to spend more time making the next article better, and less time recognizing the same wrong sentence.

Put this to work

Bring a recurring content problem. We’ll help scope the system behind it.