Home
cd ../playbooks
Creative WritingIntermediate

Clean User-Facing Text

A final text-hygiene pass for authorized prose — audits and strips invisible Unicode artifacts and homoglyphs, then reduces mechanical AI-writing signals with a measure-before/measure-after check, while preserving every fact and the writer's actual voice.

5 minutes
By guillaumemeyerSource
#editing#ai-writing-detection#unicode#proofreading#content-hygiene#technical-writing

Copy-pasted AI output can carry invisible Unicode characters that look identical to a normal space or letter but aren't — a hygiene problem entirely separate from writing style, and one that most 'sound less like AI' guides never even check for.

Who it's for: writers and editors finalizing AI-assisted articles, manuscripts, and reports they're authorized to polish, documentation teams cleaning up prose before publishing, anyone who's had AI-generated text carry stray Unicode artifacts into a final document, teams wanting an honest, non-evasion-framed approach to reducing AI writing tells

Example

"Clean up this draft before I publish it" → A deterministic pass stripping invisible Unicode and homoglyphs first, a before/after AI-cadence density check that skips the rewrite entirely if the text already reads naturally, a single rewrite pass that kills stock AI vocabulary and injects real sentence-length variation without inventing a single new fact, and an honest report distinguishing what was verifiably fixed from what's only a best-effort stylistic pass

CLAUDE.md Template

New here? 3-minute setup guide → | Already set up? Copy the template below.

# Clean User-Facing Text

A final text-hygiene pass for prose you own or are authorized to process: audit for suspicious invisible Unicode, then rewrite to reduce mechanical AI-writing signals — while preserving every fact, claim, and the writer's actual voice. Two honest boundaries up front: Unicode cleanup is deterministic; statistical-signal reduction is best-effort. Never claim a rewrite proves human authorship or is undetectable, and never use this for undisclosed-authorship evasion (academic integrity submissions, and similar).

## When to Use

Use when asked to clean, humanize, polish, or finalize articles, manuscripts, reports, documentation, emails, product copy, UI text, Markdown, or HTML prose. Don't use for code-only tasks, and don't use to help someone misrepresent AI-assisted work as undisclosed solo human authorship where that matters (an academic submission under an integrity policy, for instance).

## Workflow

### 1. Identify the Prose

Find the text readers will actually see, as distinct from code, commands, and structural markup around it.

### 2. Protect Non-Prose Spans

Before touching anything, mark as untouchable: fenced and inline code, commands, paths, URLs, identifiers, API names, exact values, formulas, citations, and anything the user explicitly asks to keep verbatim.

### 3. Preserve Every Claim

Preserve every fact, number, name, citation, and requirement exactly. **Never invent a detail, name, number, quote, or source to make the prose easier to write or more varied.** If a fact is missing, flag the gap — don't fill it. The rewrite may sharpen, compress, or reorder; it may not add or remove claims.

### 4. Measure Before Rewriting

Score the input for AI-cadence density before touching it. If the density reads as low or medium, the measurable AI-density signals are already weak — verify the text reads naturally and otherwise leave it alone. Only engage the rewrite passes when density reads as genuinely high. Rewriting text that doesn't need it is a wasted pass that risks introducing errors for no benefit.

### 5. Establish the Writing Brief

- Use a voice sample only when the user owns it or is authorized to use it — never imitate a specific named person's voice without that authorization.
- With no sample, make the prose clear and natural without pretending to imitate anyone in particular.
- Never inject a voice the source lacks: no fake first-person anecdotes, no invented specifics, no forced contrarianism, no performed candor. Preserve the writer's own deliberate rough edges and domain terms rather than scrubbing them into blandness.
- Keep any required disclosures, genuine uncertainty, and the writer's actual point of view intact.
- Pick a register that fits the text (see Voice and Domain Presets below); default to general prose when unsure.

### 6. Strip Artifacts First (Deterministic Pass)

Before any rewriting, run a mechanical pass over the text for invisible Unicode artifacts and homoglyphs — characters that look identical to a normal letter or space but aren't, often introduced silently when copying from certain AI tools or web pages. This is a clean, marker-free starting point for the rewrite, and it costs nothing to run since it's fully deterministic.

Keep whitespace semantics by default — don't collapse non-breaking spaces, narrow no-break spaces, or other meaningful whitespace unless the user specifically asks for aggressive normalization, since that can change layout in ways the writer intended.

### 7. Rewrite Once

Apply the detector-aware levers (below) in order, in a single pass:

- Vary clause order, sentence boundaries, rhythm, connectors, and function words.
- Replace formulaic transitions and filler with direct, natural wording.
- Keep the concrete details and judgment that make the text recognizably the writer's own.
- Treat unusual grammar, repetition, directness, or phrasing as possible deliberate voice or accessibility choices — change them only when asked or when they create a genuine reading problem.
- Preserve the requested language, tone, structure, and formatting; never translate unless explicitly asked.
- For non-English text, use constructions native and fluent in that language rather than English sentence patterns translated literally.
- Don't add or remove claims merely to increase variation — variation is never worth a fabricated detail.

### 8. Strip Artifacts Again

Run the deterministic Unicode pass a second time on the rewritten output, to catch anything the rewrite itself introduced — smart quotes, stray em dashes, an accidental homoglyph from a paste.

### 9. Measure After

Score the rewritten text the same way as step 4. Report scores when meaningful; a lower after-score means the measurable signals moved — it is never a verdict from any detector, and it never overrides the fact-preservation and voice rules above.

### 10. Deliver

Return the polished result. Only include an audit or explanation when explicitly asked for one.

## Detector-Aware Levers (Ordered by Effectiveness)

Statistical AI-writing detectors score probability patterns: text that's too predictable (low perplexity), too even (low burstiness), and too full of stock phrases. Only engage these when step 4's density check reads high:

1. **Strip artifacts first, always.** Invisible characters and homoglyphs are mechanical markers that hurt against every detector family and cost nothing to remove — do this unconditionally, not just at high density.
2. **Kill the stock vocabulary.** Replace AI-overused words with plain, concrete ones: *delve, tapestry, testament, underscore, foster, seamless, multifaceted, myriad, paradigm shift, harness the power of, plays a crucial role, in today's fast-paced world, it is important to note.*
3. **Inject burstiness.** Vary sentence length deliberately — a short sentence after two long ones, and occasionally the reverse. A uniform mid-length cadence is one of the strongest statistical AI signals.
4. **Flatten the structure.** Break up formulaic sections: "despite X, the future looks bright" closers, forced groups of three, "challenges and opportunities" templates, announcement-style headers.
5. **Match a real voice and keep specifics.** Prefer concrete detail from the actual source over generic phrasing. Never invent a fact to raise variance — a lower score achieved with a fabricated detail is still a failed rewrite, full stop.
6. **Normalize punctuation and assistant-voice tells.** Straight quotes, no em-dash overuse, no bolded mini-headers scattered through prose, no "I hope this helps," no "as an AI," no hedged-perfectionism disclaimers beyond what's actually required.

**Honest caveat:** these levers target statistical detectors specifically. Trained neural classifiers are adversarially trained against exactly this style of paraphrase edit, so against those, the only real lever is genuinely matching a real human distribution of writing — and even that isn't guaranteed to work.

## Voice and Domain Presets

| Preset | Personality | Rhythm | Eliminate |
|---|---|---|---|
| **General prose** (default) | Author's voice first, no injected stance | Mild variation, natural connectors | Stock AI vocabulary, uniform cadence |
| **Essay / blog** | Stance, asides, mixed feelings welcome | Strong length variation, uneven rhythm | Significance hype, aphorism formulas, rule-of-three lists |
| **Technical / documentation** | Neutral, precise | Moderate variation, short declaratives | Promotional language, em-dash overuse, bolded mini-headers; keep code and identifiers intact |
| **Academic / professional** | Formal, evidence-first | Restrained variation, controlled hedging | Over-claiming verbs, novelty padding, citation dumps; keep required discipline |
| **Business / product copy** | Plain claims, concrete value | Direct sentences | "Seamless," "empower," vague benefits, rule-of-three; keep required disclaimers |
| **Fiction** | Invented detail allowed (this genre is the exception to strict fact-preservation) | Variation to fit the narrator | Uniform cadence, editorial clichés; preserve dialect and quirks |

**Plain-language sub-mode**, for procedures, runbooks, and error messages: short common words, one instruction per sentence, imperative verbs for steps, one meaning per term, no marketing adjectives or unbounded hedging. This is a clarity floor, not a personality — it strips voice deliberately while keeping every claim and requirement intact. Use the voice-preserving presets above for essays, posts, and personal prose instead.

## Code Boundary

When prose and code are mixed, rewrite prose only. Never rename variables, alter string literals, reformat code, or change executable output as part of this pass. Preserve any executable snippet inside a Markdown or HTML file byte-for-byte wherever practical.

## Reporting

When an audit is requested, distinguish clearly between three tiers:

- **Verifiable**: Unicode characters removed or replaced, with counts; stylometry scores before and after, reported as a gauge, never a verdict.
- **Best-effort**: prose rewritten to alter token and syntax patterns.
- **Not established**: claims of detector evasion, proof of human authorship, or removal of any vendor's secret-key watermark — none of these can honestly be claimed by this process.

## Tips

- Run the artifact-stripping pass unconditionally, even on text that scores low on AI-density — invisible Unicode characters are a real, separate hygiene problem from writing style, and they're free to fix.
- The measure-before step matters as much as the rewrite itself — rewriting text that's already fine risks introducing errors or flattening a deliberate voice for zero real benefit.
- When a user provides a voice sample, confirm they actually own it or are authorized to use it before applying it — this preserves the line between "sound like my own established voice" and "impersonate someone else's."

## Limitations

- Only for text the user owns or is authorized to process — never use to help disguise AI involvement where authorship disclosure genuinely matters (a graded academic submission, a professional certification, and similar contexts).
- The statistical-detector levers are honestly limited: they work against perplexity/burstiness-style detectors, but trained neural classifiers are adversarially trained against exactly this kind of edit, and no claim here should be read as guaranteed evasion of any specific tool.
- Never a substitute for actually disclosing AI assistance where a policy, publication, or relationship requires that disclosure — this tool improves prose quality and strips accidental artifacts; it makes no authenticity claims on its own.

Get new playbooks like this one

One email a week with new Claude Code workflows. Free, like everything here.

No spam. Unsubscribe anytime.

README.md

What This Does

A two-layer text-hygiene pass distinct from a typical "sound less like AI" rewrite. Layer A is fully deterministic: strip invisible Unicode characters and homoglyphs — artifacts that look identical to a normal space or letter but aren't, and that often ride along silently when text is copy-pasted from certain tools. Layer B is a measured, honest rewrite pass: score the text's AI-cadence density first, skip the rewrite entirely if it's already low or medium, and only when density is genuinely high, kill stock AI vocabulary, inject real sentence-length variation, and flatten formulaic structure — all while a hard rule holds throughout: never invent a detail, name, number, quote, or source to make the prose easier to write or more varied. A missing fact gets flagged, never filled.

The whole approach is explicit about its own limits: Unicode cleanup is deterministic and verifiable; statistical-signal reduction is best-effort against perplexity/burstiness-style detectors specifically, and the skill refuses to claim detector evasion, proof of human authorship, or removal of any vendor's secret-key watermark. It's built for text the user owns or is authorized to process, explicitly not for disguising AI involvement where authorship disclosure genuinely matters.


Quick Start

Step 1: Create a Project Folder

mkdir clean-text && cd clean-text

Step 2: Download the Template

Click Download above, then:

mv ~/Downloads/CLAUDE.md ./

Step 3: Clean a Draft

claude

Paste in the prose you want finalized and ask Claude to clean it up. It will strip invisible artifacts, measure AI-cadence density, and rewrite only if the density check says it's actually needed — preserving every fact and the writer's own voice throughout.


Tips & Best Practices

  • Run the artifact-stripping pass unconditionally, even on text that scores low on AI-density — invisible Unicode characters are a real, separate hygiene problem from writing style, and they're free to fix.
  • The measure-before step matters as much as the rewrite itself — rewriting text that's already fine risks introducing errors or flattening a deliberate voice for zero real benefit.
  • When using a voice sample, confirm it's actually owned or authorized before applying it — this preserves the line between "sound like my own established voice" and impersonating someone else's.

Limitations

  • Only for text the user owns or is authorized to process — never use to help disguise AI involvement where authorship disclosure genuinely matters (a graded academic submission, a professional certification, and similar contexts).
  • The statistical-detector levers are honestly limited to perplexity/burstiness-style detectors — trained neural classifiers are adversarially trained against exactly this kind of edit, and no claim here should be read as guaranteed evasion of any specific tool.
  • Never a substitute for actually disclosing AI assistance where a policy or relationship requires it — this tool improves prose quality and strips accidental artifacts; it makes no authenticity claims on its own.

$Related Playbooks

Creative Writing

Chinese AI-Writing Style Audit (中文 AI 味检查)

An 11-rule checklist for stripping the tell-tale patterns of unedited AI-generated Chinese copy — em-dash chains, 不是而是 contrast pairs, 首先其次 sequence stacking, 被 passive chains — refined from real user correction feedback, each with a precise trigger, bad/good example, and scope.

5 minutes
Intermediate
Creative Writing

Novel Writing Assistant

Develop characters, plot arcs, and maintain narrative consistency across a long-form writing project.

10 minutes
Intermediate
Creative Writing

Storyboard Manager

Support creative writing with character development, story planning, chapter writing, and automated timeline tracking for narrative consistency.

10 minutes
Intermediate
Creative Writing

Voice DNA: Clone Your Writing Style

One markdown file that makes Claude stop sounding like AI and start writing in your actual voice. 95% pre-built, you just paste writing samples.

10 minutes
Beginner
Creative Writing

Worldbuilding Wiki

Keep your fictional world consistent with lore files Claude checks before creating anything new — magic systems, cultures, and history that don't contradict themselves.

10 minutes
Intermediate
Creative Writing

Worldbuilding Assistant

Create and maintain consistent fictional worlds for games, novels, or RPG campaigns with interconnected lore and rules.

10 minutes
Intermediate
Creative Writing

YouTube Script Writer

Generate complete, production-ready YouTube video scripts with hooks, structure, and pacing tailored to your channel's style and audience.

10 minutes
Beginner
Creative Writing

Book Bible — AI Novel Writing Assistant

Keep your novel consistent with an AI-powered book bible. Track characters, plot timelines, and world rules. Claude checks for inconsistencies before you write new scenes.

10 minutes
Intermediate

Browse all Creative Writing playbooks →