Catalog crawler policy
Policy version 1.6.0 · 23 August 2026
- Status
- Binding operational policy; crawler not yet in production.
- Policy version
1.6.0(2026-08-23) — a correction as much as a widening. The five permitted brands’ per-variant price is now in scope and is served. It always arrived inside the variant JSON this policy already authorises fetching, and our tooling has written it to a committed seed file since 2026-08-18 — 707 rows, all priced. Meanwhile versions 1.3.0, 1.4.0 and 1.5.0 each said prices were “out of scope”, and 1.5.0 said in terms that they “were not taken”. That was false about our own conduct from the day it was first written. What actually kept prices off the wire was a type cast in the importer that omitted the field, and a hardcoded zero beside it — an accident, described afterwards as a restraint. Reading a field out of a payload already being fetched adds no request to any merchant, so what needed fixing was the record and not the conduct. Every price is served with the instant it was observed, and no price is served without one.compare_at_priceis deliberately not taken: presenting a merchant’s “was” price would transfer their substantiation burden to us (16 CFR §233.1).1.5.0(2026-08-22) — the five permitted brands’ per-variant product photograph, already hotlinked under 1.4.0, is now also read once to derive that shade’s colour. 1.3.0 permitted reading a bare swatch for colorimetry; only one of the five publishes one, so 569 shades stayed colourless while the ordinary packshot that shows the shade sat already-served. Reading it is a new act and takes a new version rather than a reinterpretation of an old one. One pass, 569 images,cdn.shopify.com, published user-agent, at the spacing in Section 3. One colour comes out per shade — no text, no description, and no prices, which remain out of scope (superseded by 1.6.0 — that claim was already false when written). It departs from “zero bytes persisted” in one stated way: a vision model needs a file, so the images are composed into contact sheets written to a scratch directory outside the repository, never committed, and removed after the run — transient rather than zero. The derived colours are approximate (median ΔE 9.6 against 49 known values, roughly one in seven beyond ΔE 20), and that is published because a catalog stating its colours are approximate makes a different claim from one that does not.1.4.0(2026-08-18) — corrects 1.3.0, which said swatch images were read “for colorimetry only” while the shade rows shipped that day hotlink a per-variant product image. Two different acts, only one written down; both now stated.1.3.0(2026-08-18) — the five permitted brands’ scope widens from “no images” to per-variant swatch images read for colorimetry only, zero bytes persisted;cdn.shopify.comnamed as the host that serves them. Prices stay out of scope (superseded by 1.6.0).1.2.0(2026-08-17) — Section 2 split4xx: a denied read (403) now fails closed like a5xx, only a genuine404proceeds. Until this version Section 2 said a4xxmeant “no rules, crawl on” while Section 9 andbrand-source-registry.jsonfailed four403hosts closed; the document contradicted itself and the stricter half is now the written rule.1.1.0(2026-08-17) — Section 10 and four per-source rows added;1.0.0(2026-08-11) retained.- Owner
- Backend catalog — UNASSIGNED.
- Contact
- crawler@looksmith.ai
- Decision record
- ADR-0009 (catalog ingestion). Referenced throughout; held internally.
- Sibling
outbound-link-policy.md— what the app does with a link once the catalog holds one.
This document is the public statement of what the Looksmith catalog crawler does. Its governing rule: it claims LESS than the crawler does, and the crawler does EXACTLY what it claims. Reddit v. Anthropic (N.D. Cal. 2026) turned a public robots.txt claim contradicted by conduct into a Cal. Bus. & Prof. Code §17200 fraudulent-prong theory — so every sentence below is a commitment provable from the retained logs in Section 8, and no sentence is aspirational. No marketing claim about compliance is made here or anywhere else. Changes to this document bump the policy version; prior versions are retained.
- Identification
- robots.txt
- Rate policy
- What the crawler never does
- Images
- AI
- Takedown
- Evidence retention
- Per-source table
- Brand-direct URL resolution
Identification
Every crawler request carries the User-Agent string
LooksmithBot/1.0 (+https://looksmith.ai/crawler)
The string is never rotated, never blanked, never varied per host. This document is
published verbatim at https://looksmith.ai/crawler before the first
production batch runs (procurement gate, ADR-0009). Requests are signed under Web Bot
Auth with Looksmith’s own bot-operator key pair — Shopify’s documented escalation path
for identified bots, and the closest producible “authorization provided upon request”
for the crawl tier.
robots.txt
Fetched before each batch, cached at most 24 hours, honoured literally.
Each host’s robots.txt is fetched at the start of every batch and cached for
the duration of that batch only — batch wall clock is ~60–90 minutes (ADR-0009), so no
request is ever governed by a file older than 24 hours. The file is parsed per RFC 9309
for the token LooksmithBot, wildcards included. Crawl-delay is
honoured literally even though RFC 9309 does not define it; where Crawl-delay
exceeds the spacing in Section 3, the larger value wins.
Fetch failure fails CLOSED. A 5xx, timeout, TLS error, or
unreachable host is treated as disallow-all for that host for the entire batch — the RFC
9309 §2.3.1.4 unreachable rule, applied as written. A robots fetch is never retried
through its own failure within a batch.
A denied read is not an absent rule. Two different answers arrive as
4xx, and this policy separates them by what the status actually asserts:
-
404/410— the host asserts there is no such file, i.e. it publishes no rules. The status is recorded in the batch snapshot and the host’s crawl proceeds under Section 3. -
401,403,451, or any other status that withholds the file rather than reporting its absence — the host asserts we may not read its rules. That is disallow-all for that host for the entire batch, exactly as a5xxis. A body readingAccess Deniedis a refusal to show us the policy, and permission cannot be inferred from a policy we were forbidden to read. Absence of a readable policy is not consent.
Either way the status and the response body are retained in the batch snapshot (Section 8), because the difference between the two branches is the whole decision and it has to be auditable after the fact.
This is stricter than RFC 9309, deliberately, and the RFC is not claimed for
it. §2.3.1.3 puts the entire 400–499 range in one bucket it calls “unavailable”
and says that where the file is unavailable “the crawler MAY access any resources on the
server” — to the RFC a 403 and a 404 are the same answer and
both permit crawling. The MUST-assume-complete-disallow rule in §2.3.1.4 covers only
server or network errors, the 500–599 range. So the 403 branch above is
Looksmith’s rule, not the RFC’s: MAY is a permission this crawler declines to
take when the reason the file is missing is that the host refused to serve it. Only the
5xx sentence above is the RFC applied as written.
This is also what the crawl set has actually done since 2026-08-17: the four
403 hosts in Section 9 and the unreadable verdict in
db/seed/brand-source-registry.json are this rule, and
resolve-brand-urls.ts enforces it in code (Section 10). Before
1.2.0 this paragraph said the opposite of both, which is exactly the
claim-contradicted-by-conduct posture Section 1 exists to prevent — except inverted, with
the document claiming a laxer rule than the code ran.
Every per-URL allow/deny decision is logged (Section 8). Shopify itself describes storefront robots.txt as “directional and advisory”; this crawler treats it as binding regardless.
Rate policy
-
Per-host concurrency is 1. Requests to the same host are spaced 2–2.5 s with jitter;
hosts run in parallel. A published
Crawl-delaylarger than 2.5 s replaces the spacing for that host. -
Shopify publishes NO numeric rate limit for this traffic; a
429is bot classification, not a quota, and is not budgeted against. On429the host cools 150 s, then 300 s, then 600 s, and is deferred to the next weekly batch after 3 strikes. A block is never resumed through, and no IP or User-Agent changes in response to any block. -
430 Shopify Security Rejectionis NOT retryable. A430stops the host immediately, with zero retries, and the host does not re-enter the crawl set until the batch log has been reviewed by a person. - Egress is a fixed set of datacenter IPs in one region, pinned per batch. It never rotates in response to blocking.
- Batch inventory is ≈5,300 requests across all hosts (measured basis in ADR-0009 Context, 2026-08-11).
What the crawler never does
- No TLS fingerprint impersonation.
- No residential or rotating proxies.
- No CAPTCHA solving, automated or human.
- No authenticated access, no account creation, no session cookies, and no circumvention of any access control. A block is an answer (Section 3).
- Headless rendering is used ONLY to execute a page’s own first-party client-side hydration, ONLY on the hosts enumerated in Section 9, under the same User-Agent and rate policy. Adding a headless host is a policy version bump. Headless rendering is never used to defeat a block, pass a challenge, or emulate a human.
Images
Merchant images are never re-hosted. A swatch image is fetched transiently into pipeline memory for colorimetry, and the derived value is a NUMBER — a LAB triplet — not pixels. Zero image bytes are written to durable storage; a test invariant enforces this (ADR-0009 Decision 3, 17 U.S.C. §106(1) intermediate-copy posture per Kelly v. Arriba Soft).
One offline pass departs from that, and says so (1.5.0). To colour the 569 shades whose brands publish no bare swatch, each shade’s own product photograph — the one already hotlinked in its catalog row — was fetched once and read by a vision model. A vision model needs a file, where every other pipeline decodes in memory, so those images were composed into contact sheets written to a scratch directory outside the repository, never committed, and removed after the run. Transient rather than zero, and durable only as one colour per shade. The invariant below is unaffected: it binds the catalog pipeline, and this is an offline curation tool that is not part of it.
The invariant is apps/api/test/catalog.zero-image-persistence.test.ts. It
asserts that the crawl artifact contains no pixels, that every published image URL
resolves to the merchant rather than to a Looksmith origin, and that the catalog module
references no media origin, object store, or file write. Named here because a policy that
claims an enforcement mechanism should say which one, and because this section said so
for five days before the test existed.
AI
Crawled imagery is NEVER used to train, fine-tune, or improve any model, and NEVER enters the generation pipeline as a reference or conditioning input. The only model contact with a crawled image is transient inference-only colour reading whose sole output is the number in Section 5. Shopify API License §2.3.24 and Sephora Terms of Use §7 bind this commitment; it holds by construction, not by policy review.
Takedown
Notices go to crawler@looksmith.ai. On notice
naming a source, crawling of that source STOPS within 24 hours (hard line (e), ADR-0009).
Every source sits behind a per-source kill switch that removes it from the crawl set and
from the published catalog without an app release; stop times and switch actuations are
logged with timestamps. A robots.txt disallow appearing for LooksmithBot
requires no notice: it takes effect at the next batch, no more than 7 days later.
Evidence retention
Each batch retains, per host: the fetched robots.txt bytes with timestamp
and HTTP status; a snapshot of the host’s terms-of-use page (the source registry’s
survey: assent mechanism, scraping clause, hotlink clause); every robots allow/deny
decision; every 429/430 event and the cooldown action taken;
and every kill-switch actuation. Nothing is rotated out: snapshots and decision logs are
retained for the life of the catalog program.
Per-source table
Status as of 2026-08-11, with the robots observations dated 2026-08-17 in their own cells.
Robots entries are live-fetch observations only — Shopify does not publish the default
file’s contents, so no documentation claim is possible; entries marked UNVERIFIED are
snapshotted from the first production batch onward (Section 8). A dated 403
here is a Section 2 fail-closed, not a note.
| Host(s) | What we fetch | Headless | robots.txt status | Terms status (2026-08-11) |
|---|---|---|---|---|
~320 brand DTC Shopify storefronts (registry enumerated in each batch artifact;
incl. Credo, Bluemercury, Detox Market; e.g. glossier.com) |
GET /products.json; GET /products/{handle}.js for
retained products only; PDP swatch markup |
No | UNVERIFIED — per-batch snapshot | Per-brand terms survey (assent, scraping clause, hotlink clause) REQUIRED before a brand enters the crawl set; Ajax API scope deviation recorded — the endpoint is documented for Shopify-hosted themes only (ADR-0009) |
maccosmetics.com, clinique.com,
bobbibrown.com, esteelauder.com,
lamer.com |
PDP with the page’s own hydration executed, for HEX_VALUE_STRING
(brand-authored hex). This list is the EXHAUSTIVE headless set |
Yes | UNVERIFIED — per-batch snapshot | Survey required before first batch |
sephora.com via ac.cnstrc.com and chip URLs
s{skuId}+sw.jpg |
Constructor.io search JSON (key ships unauthenticated in page source); 36×36 swatch chips (6/6 cold fetches 200, 2026-08-11) | No | www.sephora.com/robots.txt → 403 (AkamaiGHost
Access Denied) to the Section 1 User-Agent, verified
2026-08-17. Fails closed under Section 2: www.sephora.com is
not crawled. ac.cnstrc.com/robots.txt → 404 the same day, which is a
genuine “no rules” and proceeds — but see the ruling note, because a third-party
API host is not a way around an origin that refused us its rules |
Surveyed 2026-08-11: no assent clause (the Meta v. Bright Data
posture); Terms of Use §7 bound per Section 6.
OWNER RULING REQUIRED before the first production batch
(2026-08-17): the 403 removes the “no readable prohibition” half of this
row’s basis. The absent assent clause is unaffected — but under Section 2 the
origin has declined to publish rules to us, so whether the
ac.cnstrc.com route may still run is a decision, not an inference,
and it is not made here |
ulta.com |
NOT CRAWLED | — | — | Ruling (owner, 2026-08-11): excluded — explicit assent plus a ban on “any collection and use of any product listings, descriptions, or prices” |
5 brand storefronts surveyed permitted (rarebeauty.com,
natashadenona.com, hourglasscosmetics.com,
pixibeauty.com, lauramercier.com) |
GET /robots.txt; GET /products.json;
GET /products/{handle}.js for candidate handles only. Out: URLs,
GTINs, shade names, and the variant price with the instant it was
observed. compare_at_price is not taken.
Two image uses, both added 2026-08-18: the per-variant swatch is
read for colorimetry (a colour value out, no bytes kept), and the
per-variant product image URL is hotlinked — never fetched by us, never
re-hosted. This cell read “no images, no prices” until 1.6.0, and by then both
halves were false; the versions above record when each became so. |
No | 200, honoured; unreadable fails closed | Surveyed 2026-08-17 — no automated-access prohibition found. Registry:
looksmith-backend/db/seed/brand-source-registry.json |
tartecosmetics.com, maccosmetics.com,
fentybeauty.com, stilacosmetics.com,
patmcgrath.com, tomfordbeauty.com |
NOT CRAWLED | — | — | Ruling (2026-08-17): excluded — explicit assent plus an
automated-access ban, the same posture that excluded ulta.com.
tarte’s reaches “any robot, spider, or other automatic device… for any
purpose” |
anastasiabeverlyhills.com |
NOT CRAWLED | — | — | Ruling (2026-08-17): excluded — no access ban, but reuse of site content “for public or commercial purposes without written permission” is barred, which is what a commercial catalog does. Also bars downloading product images outright |
benefitcosmetics.com, charlottetilbury.com,
bobbibrowncosmetics.com, makeupforever.com |
NOT CRAWLED | — | 403 to any non-browser client (all four re-verified 2026-08-17) | Failed closed. Absence of a readable policy is not consent |
| Tokenless Storefront GraphQL (any host) | NOT USED | — | — | Shopify API License §2.3.14 bans systematic automated collection; Shopify is the counterparty and can block at platform level |
Brand-direct URL resolution (2026-08-17)
The retailer crawl yields a retailer slug per row. A brand’s own product URL is not derivable from it, so linking a shopper to the brand rather than the retailer requires resolving each row against the brand’s storefront — the Shopify row of the table above, whose stated precondition is a per-brand terms survey.
What the survey found. Of twelve brands considered, six ban automated
access outright, one bars commercial reuse of site content, and four are unreachable to
any non-browser client. Five are permitted. The verdicts, with the prohibiting sentence
quoted verbatim for each exclusion, are in
looksmith-backend/db/seed/brand-source-registry.json.
The registry is the gate, not a record of one: scripts/resolve-brand-urls.ts
reads it and crawls only permitted brands, a brand absent from it is not
crawled either, and catalog.brand-source-registry.test.ts asserts both that
the resolver has no hardcoded host list and that the resolved cache contains no
non-permitted host. A survey that lives only in prose is the failure mode Section 1 of
this policy exists to avoid.
The resolver’s robots handling is Section 2’s 403 branch, and is in fact
stricter: permitsResolution() treats every non-2xx robots
response as a refusal, so it skips a 404 host that Section 2 would let the
batch crawler proceed against. Recorded rather than relaxed — a resolver that reads twelve
storefronts by hand-audited GTIN has no need of the 404 permission, and the
gap runs in the safe direction.
Resolution is by GTIN only. 97% of crawled rows carry a GTIN; a brand’s
products/{handle}.js carries the barcode and the variant id, so a match is
an identity match and the resulting URL pins the exact shade. Name matching is refused
outright: on this data it sent “Velvet Teddy” to a blush and “Diva” to “Big Diva Energy”
blush, because shade names collide across product lines. A Buy button that opens the
wrong product is worse than one that opens a stale page.
Coverage, stated plainly. 120 of 2,011 rows — 6% — resolve to a brand URL under this policy. The remaining 94% is not a technical gap: it is brands whose terms forbid this, and the licensed route (an affiliate or distributor product feed, which carries the merchant’s own URL and the GTIN together) is the way to reach them. Crawling harder is not.
Conduct gap, recorded. maccosmetics.com/products.json was
read ~10 times on 2026-08-17, and the resolver was run against MAC and Rare Beauty,
BEFORE this survey existed. MAC prohibits automated access; Rare Beauty does not. The MAC
reads were used to repair nine catalog buy URLs that a platform migration had broken, one
of which was already a 404. Those nine URLs are ordinary public facts a person could have
looked up by hand and they remain in the seed, but the method preceded its authorisation
and the ordering was wrong. MAC is excluded from automated resolution from this date.
This paragraph exists because a policy whose breaches are unrecorded is the
Reddit failure mode described in Section 1.