How to automate Shopify product descriptions at catalogue scale
Alex
Automating Shopify product descriptions means generating the copy for each product from your own structured product data, instead of writing every one by hand or pasting the supplier's text. It is worth doing above roughly a hundred SKUs, where writing by hand stops being possible and pasting stops being safe.
The order matters more than the tool. Clean data first, then a template that separates what is fixed from what varies, then a review pass weighted toward the products that make the money, then a measurement that compares treated products with untreated ones rather than the store before and after.
Why automate product descriptions at all?
Two reasons, and only one of them is about time.
The time argument is obvious: a thousand SKUs at fifteen minutes each is 250 hours of writing, which nobody has. The second reason is the one that actually costs money. The supplier description on your product page is on every other store selling that product. Google then has to choose one version to show, and there is no reason for it to choose yours.
Is duplicate content a Google penalty?
No, and the distinction matters for what you do about it. Google does not apply a duplicate content penalty to a store that reuses a manufacturer's text. What it does is deduplicate: it picks one version as canonical and the others simply do not appear in results. The outcome feels like a penalty and has a different cure.
A penalty would mean you did something wrong and must undo it. Deduplication means you offered nothing distinguishable and must add something. That is why "spinning" the supplier text into synonyms fails: the page is still the same page with different adjectives. What differentiates a product page is information the supplier did not write, which is where the next section starts.
What changes above a few hundred SKUs?
The economics invert. Below about a hundred products, hand-writing wins: it is better, and the total cost is bearable. Above a thousand, hand-writing is not a slower option, it is not an option at all, and the real choice is between generated copy and no copy.
The middle is where the decision is genuinely open, and the deciding factor is not the SKU count but the rate of change. A catalogue that renews every season needs a repeatable process regardless of size, because the writing is never finished.
What has to be in your product data first?
Generation quality is bounded by input quality, and this is the step everyone shortens. A model given a title and a price will write something plausible and empty, because there was nothing else to say. The fix is upstream and it is unglamorous.
Which fields need to exist before you generate anything?
| Field | Why the generation needs it | What happens without it |
|---|---|---|
| Product type and category | Decides the angle: a technical spec sheet or a lifestyle pitch | Generic copy that suits neither |
| Materials, composition, dimensions | The concrete detail a buyer looks for and a supplier text usually buries | Adjectives instead of facts, which convince nobody |
| Variants (size, colour, capacity) | Avoids five near-identical pages competing with each other | Internal cannibalisation across your own variants |
| Use case or target buyer | Turns a feature into a benefit, which is what actually sells | A spec list with no reason to buy |
| What makes it different from the neighbouring product | The only genuinely non-duplicable input | The exact copy your competitor also generated |
Note what is not on the list: keyword volume, meta-title length, any SEO variable. Those come later and they are mechanical. This table is about having something to say.
What do you do with the products whose data is incomplete?
Three options, and choosing per product beats applying one rule to the whole catalogue.
- Enrich by hand, for the products that carry the revenue. It is a short list and it is worth the hours.
- Generate on a reduced angle, for the long tail: honest, shorter copy built only on the fields that exist. Better than a padded page pretending to specifics it does not have.
- Skip, for products about to be discontinued. Writing copy for stock that leaves in six weeks is a pure loss.
The failure mode worth naming: instructing the generator to "be creative when data is missing". It will be, and it will invent a composition or a dimension. On a product page, an invented specification is not an SEO problem, it is a returns and consumer-law problem.
How do you generate copy that is not generic?
What is fixed in the template and what varies per product?
A workable template separates three layers. The brand layer is fixed: tone, address form, what you never say, the sentence about shipping or returns. The category layer varies by product type: a running shoe and a coffee grinder do not deserve the same plan. The product layer is the injected data itself.
Getting this wrong in the easy direction gives 2,000 pages with the same skeleton, recognisable at a glance and useless for differentiation. Getting it wrong in the other direction gives 2,000 pages with 2,000 personalities, which reads as an incoherent store.
How do you keep 2,000 descriptions from saying the same thing?
The lever is not asking for variety in the instructions, it is varying the input. Two products that go in with the same five fields come out with the same description, whatever tone you asked for. The differentiating field, "what makes this one different from its neighbour", is what breaks the tie.
The check is cheap and worth automating: extract the first sentence of every generated description and count the duplicates. If the same opening appears forty times, the template is doing the work instead of the data.
What actually ranks on a product page?
Rarely the description alone. A product page ranks on the whole thing: the title, the structured data, the images, the reviews, the internal links pointing at it. The description matters because it is the part that carries the specifics and answers what the buyer is actually wondering.
Which is why the format that works is the same one that works for answer engines: the question a buyer would ask, answered directly. "Is it machine washable?" answered in one line beats a paragraph of atmosphere. It also gives you a real chance in the assistant answers where people increasingly compare products before buying.
Where do collections and internal links fit?
They are the part of the job an isolated description cannot do, and the part that is easiest to automate well because it is mechanical.
A generated description can carry two kinds of link: up to the collection the product belongs to, and across to complementary or alternative products. The first helps the collection page accumulate the signals it needs to rank on the category query, which is where the volume is. The second keeps a visitor who was on the wrong product inside the store.
Two guardrails. Insert links only where they are genuinely relevant, because ten links at the bottom of every product page is a footer, and it is treated as one. And generate them from the catalogue structure rather than by hand, so a discontinued product does not leave dead links across three hundred pages: on Shopify a removed product returns a 404, and nothing warns you that other product pages still point at it.
Collection pages deserve their own treatment. A collection with a real introduction, answering what the category is and how to choose within it, is a page that can rank on a query with far more volume than any single product. Most stores leave it as a bare grid of thumbnails.
How do you handle a multilingual catalogue?
Generate in each language, do not translate. A translated description carries the source language's assumptions: units, sizing conventions, seasonality, the arguments that convince in that market. A description generated directly in the target language, from the same structured data, starts from the buyer in that country.
The mechanics on Shopify are the standard ones and they are not optional: one URL per language, hreflang pointing both ways between the pairs, and a canonical on each version pointing at itself. A one-way hreflang is ignored, which is the most common way a multilingual store quietly gets nothing for the work.
What review checkpoints keep quality up?
Reviewing everything is not realistic and reviewing nothing is how a catalogue ends up with an invented specification on a page nobody looked at. Weight the effort instead.
| Tier | Which products | Review |
|---|---|---|
| Hero | Best sellers and high margin, roughly the top 5 % | Read and edit every one, by hand |
| Middle | Products with real traffic | Sample, roughly one in ten, plus every automated check |
| Long tail | The rest | Automated checks only, spot-check on the errors they raise |
The automated checks that catch most of it: does the description mention a specification absent from the product data (the hallucination test), does it repeat the first sentence of another product, does it contain a price or a delivery promise the store cannot honour, and is it within the length you decided.
Say the honest part out loud: generated copy is a first draft that gets published, and the review is what makes that acceptable. Any supplier claiming zero human involvement is describing a risk, not a feature.
How do you measure whether it worked?
The instinct is to compare the store before and after. It is the one method that cannot work: seasonality, prices, campaigns and stock all moved at the same time, and the description change is buried under them.
Compare treated products with untreated ones over the same period. Rewrite half the catalogue, leave the other half as it is, and read the difference. It is the only reading that survives a busy quarter.
| Metric | What it tells you | How long before it means anything |
|---|---|---|
| Impressions on product pages | The page has become eligible on more queries | 2 to 4 weeks, and it moves first |
| Average position | The trend, never a single number | 4 to 8 weeks |
| Organic clicks | The real objective | Slow, and noisy on a small cohort |
| Conversion rate on the page | Whether the copy sells, not just whether it ranks | Needs volume before it means anything |
Three measurement traps, all measured on properties we run rather than repeated from elsewhere. Search Console anonymises rare queries: on one property the reported total was 892 clicks while the per-query rows summed to 560, a 37 % gap, so a total rebuilt by summing rows is wrong by construction and long-tail product queries are exactly what gets hidden. Average position is weighted by impressions, so it does not average across pages and only compares between windows of the same width. And finalised data lags about two days, so a baseline taken on same-day figures is not reproducible.
Last one, the least intuitive: check that the pages have actually been recrawled before concluding anything. Search Console's URL Inspection API returns the last crawl time per URL. Without it, "the rewrite did nothing" and "nothing has looked at the rewrite yet" are indistinguishable, and they call for opposite decisions.
Where does Wisewand fit?
Wisewand connects to Shopify natively, through an extension published on the Shopify App Store, so generated content is published into the store rather than exported and pasted back.
It has two modes. In Autopilot, it picks and schedules the topics of a rolling 30-day editorial calendar itself, writes about one article a day and publishes it automatically, at 77 EUR per month after a 3-day trial for 1 EUR. In Manual mode you drive: one credit is worth roughly one article, credits stack up and never expire, and the content types live there, which is where product pages and category pages belong. Content can be produced in around 107 languages, which is what makes generating per market realistic instead of translating.
What no tool can promise, and it is worth stating rather than implying: a position on Google or a citation in an AI answer. What is controllable is producing pages that carry real specifics, are linked to their collections and are actually reviewed. The same discipline applied to a blog rather than a catalogue is covered in our guide to getting cited by AI answer engines.
FAQ
How do I prepare my Shopify product data before automating descriptions?
Make sure five things exist per product: type and category, materials and dimensions, variants, the use case or target buyer, and what makes this product different from the one next to it. That last field is the only genuinely non-duplicable input. A generation given only a title and a price produces something plausible and empty, because there was nothing else to say.
Will Google penalise me for supplier product descriptions?
There is no duplicate content penalty. Google deduplicates: it picks one version as canonical and the others simply do not appear. The practical outcome resembles a penalty but the cure is different, which is why rewording the supplier text with synonyms does not work. The page has to carry information the supplier did not write.
How many SKUs make automation worth it?
Roughly a hundred is where hand-writing stops being realistic. Below that, writing by hand wins on quality at a bearable cost. Above a thousand, the real choice is between generated copy and no copy at all. In the middle, the deciding factor is not the count but the rate of catalogue renewal: a collection that turns over every season needs a repeatable process whatever its size.
How do I stop 2,000 generated descriptions from sounding identical?
Vary the input, not the instructions. Two products entering with the same five fields leave with the same description whatever tone was requested. Add a differentiating field per product, then run a cheap check: extract the first sentence of every description and count duplicates. If one opening appears forty times, the template is doing the work instead of the data.
How do collections and internal links fit into this?
A generated description can link up to its collection and across to complementary products. The first helps the collection page rank on the category query, where the volume is; the second retains a visitor who landed on the wrong product. Generate the links from the catalogue structure, because a discontinued product returns a 404 on Shopify and nothing warns you that other pages still point at it.
Should I translate my descriptions or generate them per language?
Generate in each language. A translation carries the source market's units, sizing and arguments. Generating from the same structured data starts from the local buyer instead. On the technical side: one URL per language, hreflang declared in both directions, and a self-referencing canonical. A one-way hreflang is ignored, which is the most common way this work quietly yields nothing.
Which descriptions should a human actually review?
Weight the effort. Read and edit every hero product (best sellers and high margin, about 5 % of the catalogue), sample roughly one in ten of the products with real traffic, and rely on automated checks for the long tail. The checks worth running: specifications absent from the product data, repeated opening sentences, prices or delivery promises the store cannot honour.
Can the automation invent product specifications?
Yes, if you instruct it to fill gaps creatively. On a product page an invented composition or dimension is not an SEO problem, it is a returns and consumer-law problem. Instruct the generation to write only from the fields provided and to stay shorter when data is missing, then run a check comparing every specification mentioned against the product record.
How long before automated descriptions show a result?
Impressions on product pages move first, typically in two to four weeks; average position follows over four to eight; clicks are slower and noisy on a small cohort. Before concluding anything, check the pages were recrawled using Search Console's URL Inspection API. Otherwise "the rewrite did nothing" and "nothing has looked at it yet" are indistinguishable.
How should I measure the impact without fooling myself?
Do not compare the store before and after: seasonality, pricing and campaigns all moved too. Rewrite half the catalogue and compare it with the untreated half over the same period. Also note that Search Console anonymises rare queries (we measured a 37 % gap between the reported total and the sum of per-query rows), and long-tail product queries are precisely what gets hidden.
Does Wisewand publish directly into Shopify?
Yes. The connection is native, through an extension published on the Shopify App Store, so content is published into the store rather than exported and pasted back. In Autopilot mode publishing is automatic; in Manual mode you keep control of what goes out and when.
This article also exists in French: le lire en français