Product Descriptions at Scale: 5,000 SKUs, One Brand Voice, Zero Copy-Paste Detection
47% of online sellers now use AI for product descriptions — and most of them produce the same interchangeable slop, or worse, still run manufacturer copy that five competitors also publish. Here's the pipeline that actually works at catalog scale: a codified voice, structured data in, template variation, per-SKU uniqueness from real attributes, and the QA gates that catch the AI's lies.
- ecommerce
- content
- aeo
Every catalog past a few hundred SKUs hits the same wall: you can't afford to hand-write every description, so you either recycle manufacturer copy (which five competitors also publish — and if five retailers use the same description, none of them win) or you leave products with thin, generic filler that ranks for nothing and convinces nobody. AI supposedly solved this — 47% of online sellers now use AI for product descriptions, with time savings around 88% versus manual writing — but walk through most AI-written catalogs and you find the third failure mode: five thousand technically-unique pages that all sound like the same beige robot, and increasingly like every other store's beige robot too. This post is the pipeline that avoids all three fates: one brand voice, genuinely differentiated copy per SKU, at a scale no content team could hand-write.
First, a reframe on the title's last clause. "Zero copy-paste detection" doesn't mean tricking AI detectors — that's a losing game and the wrong goal (and Google has been explicit that it rewards helpful content regardless of how it's produced). It means passing the two detectors that actually matter: the duplicate-content problem (your copy matching the manufacturer's, your competitors', or your own other SKUs so closely that search engines see nothing worth indexing) and the human sameness detector (a shopper reading three of your product pages and feeling the copy was stamped, not written). Both are solved the same way: descriptions generated from unique structured data about each product, not around it.
Why naive AI generation fails at SKU 500
The default approach — paste the product title into a generator, publish the output — collapses at scale for reasons that are predictable in hindsight:
Same prompt, same shape. Ask a model for 500 descriptions with the same instruction and you get 500 outputs with the same rhythm: an adjective-rich opener, three feature bullets, a "perfect for" closer. Technically unique text, structurally identical slop. Shoppers notice; so do the AI engines now reading your catalog for shopping answers, which reward attribute-dense specificity and skip interchangeable copy.
No data in, hallucination out. Given only a title, the model invents materials, dimensions, and benefits. At 5,000 SKUs, even a 2% hallucination rate is 100 product pages that lie — and in e-commerce, inaccurate descriptions come home as returns, disputes, and marketplace suppressions. The fidelity rule from our AI-imagery post applies verbatim to text: the description is a legal representation of what the buyer receives.
Brand voice by adjective doesn't survive scale. "Write it in a premium, friendly tone" produces a different interpretation of "premium and friendly" every hundred generations. Voice drift is invisible per-SKU and glaring per-catalog.
The fix isn't a better one-shot prompt. It's a pipeline.
The pipeline: five stages from data to published copy
Stage 1 — Structured data first (the stage everyone skips). A description generator is only as good as its inputs, and the input should be a per-SKU attribute record: materials, dimensions, care, fit, compatibility, use cases, what's in the box, and — the differentiators AI can't invent — your proof points: review themes ("runs small, size up" appears in 40 reviews), bestseller status, origin story, warranty terms. On Shopify this lives naturally in metafields; elsewhere, a PIM or even a disciplined sheet works. This stage is why the output ends up unique per SKU: uniqueness comes from the data, not from asking the model to "be creative." Two products with genuinely different attribute records cannot produce the same description; two products with only a title each almost certainly will.
Stage 2 — Codify the voice as rules and exemplars, not adjectives. A working brand-voice spec is concrete: 8–10 of your best existing descriptions as exemplars, explicit do/don't vocabulary (say "handloomed," never "artisanal"; never open with "Introducing"), sentence-length and formality parameters, how you handle claims ("cite the review count, never say 'customers love'"), and per-category register shifts (bedsheets ≠ electronics). This spec is a version-controlled asset — the same philosophy as the skill files we build for our own production workflows: prompting as configuration, reviewed and improved over time, not re-improvised per batch.
Stage 3 — Template variation by design. Structural sameness is defeated deliberately: 4–6 description architectures per category (problem-first, sensory-first, spec-led, use-case-led, story-led), assigned across the catalog so adjacent products in a collection never share a shape. Add per-category required elements (fit guidance for apparel, compatibility for electronics, care for textiles) and length variance. The catalog reads as written by one person on different days — which is exactly the target.
Stage 4 — Generate in graded tiers. Not all SKUs deserve equal investment, and the successful pattern in every case study is triage: Tier A (top ~10% by revenue/traffic) gets generated drafts plus real human editing and enrichment; Tier B (the middle) gets full pipeline generation with human spot-review; Tier C (long tail) gets pipeline generation with automated QA only. Start with a pilot batch of 25–50 diverse SKUs, review hard, tune the voice spec and templates, then scale — the refinement phase is what prevents generating five thousand pages that all need revision.
Stage 5 — QA gates before anything publishes. The same discipline as our AI-QA post, applied to copy: automated checks for factual fidelity (every claim in the output must trace to an attribute in the input record — flag anything unsupported), banned-phrase and voice-compliance linting, cross-SKU similarity scoring (embedding similarity between descriptions; anything above threshold gets regenerated), length/format validation, and SEO basics (target term present naturally, no keyword stuffing). Plus the sampling loop: a human reads a random 5% of every batch, and every caught failure becomes a new automated check.
The SEO and AEO layer (because these pages now serve two readers)
Each description now has two audiences: the shopper, and the AI engines composing shopping answers. Serving both is mostly the same work: open with a direct, factual statement of what the product is and who it's for (the extraction zone); include a usage-scenarios block — "fits carry-on requirements," "suitable for apartments up to 60m²" — which is what matches "what should I buy for..." queries; keep specs in a structured table, mirrored in Product schema rather than buried in prose; and write genuinely distinct copy per variant where variants have their own pages, or consolidate with canonicals where they don't. One measurable upside worth chasing: brands doing this properly report meaningful conversion lifts (studies cite up to ~23% for optimized AI-written catalogs versus generic copy) — because specific, accurate, scenario-rich descriptions answer the questions that stall purchases.
And the multilingual multiplier, which matters enormously for Indian and international brands: the same pipeline extends to translation-first publishing — attribute records are language-neutral, so generating Hindi, Gujarati, or Arabic descriptions from the same data (with a per-language voice spec, not machine-translated English) turns one catalog investment into several storefronts. We've done exactly this for trilingual product pages, and the pipeline cost of language two is a fraction of language one.
What this looks like operationally
For a 5,000-SKU catalog, the realistic shape of the project: 1–2 weeks building the attribute layer and voice spec (the unglamorous majority of the value), a pilot week of tuning on 50 SKUs, then batch generation at a few hundred SKUs per day through the QA gates — with Tier A human editing running in parallel. Ongoing, the system earns its keep on change: new arrivals described on ingestion, seasonal refreshes regenerated from updated templates, and review-theme enrichment feeding back into descriptions quarterly (the "runs small" insight belongs in the copy, and no manual process ever puts it there).
The strategic point under all the mechanics: at 5,000 SKUs, your product copy is infrastructure, not writing. Treated as writing, it decays into some mix of duplicate, thin, and beige. Treated as infrastructure — structured data, versioned voice, templated variation, automated QA — it becomes the rare asset that improves search visibility, AI-shopping visibility, and conversion simultaneously, and keeps improving as your data does.
We build this pipeline end to end — the metafield/PIM attribute layer, the voice spec, the generation and QA automation, and the multilingual extension — usually as a 4–6 week engagement that pays for itself in the first season. If your catalog is running on manufacturer copy or first-generation AI slop, the gap between you and this system is currently your competitor's opportunity.