LOOKSMITH

Catalog crawler policy

Policy version 1.6.0 · 23 August 2026

Status
Binding operational policy; crawler not yet in production.
Policy version
1.6.0 (2026-08-23) — a correction as much as a widening. The five permitted brands’ per-variant price is now in scope and is served. It always arrived inside the variant JSON this policy already authorises fetching, and our tooling has written it to a committed seed file since 2026-08-18 — 707 rows, all priced. Meanwhile versions 1.3.0, 1.4.0 and 1.5.0 each said prices were “out of scope”, and 1.5.0 said in terms that they “were not taken”. That was false about our own conduct from the day it was first written. What actually kept prices off the wire was a type cast in the importer that omitted the field, and a hardcoded zero beside it — an accident, described afterwards as a restraint. Reading a field out of a payload already being fetched adds no request to any merchant, so what needed fixing was the record and not the conduct. Every price is served with the instant it was observed, and no price is served without one. compare_at_price is deliberately not taken: presenting a merchant’s “was” price would transfer their substantiation burden to us (16 CFR §233.1). 1.5.0 (2026-08-22) — the five permitted brands’ per-variant product photograph, already hotlinked under 1.4.0, is now also read once to derive that shade’s colour. 1.3.0 permitted reading a bare swatch for colorimetry; only one of the five publishes one, so 569 shades stayed colourless while the ordinary packshot that shows the shade sat already-served. Reading it is a new act and takes a new version rather than a reinterpretation of an old one. One pass, 569 images, cdn.shopify.com, published user-agent, at the spacing in Section 3. One colour comes out per shade — no text, no description, and no prices, which remain out of scope (superseded by 1.6.0 — that claim was already false when written). It departs from “zero bytes persisted” in one stated way: a vision model needs a file, so the images are composed into contact sheets written to a scratch directory outside the repository, never committed, and removed after the run — transient rather than zero. The derived colours are approximate (median ΔE 9.6 against 49 known values, roughly one in seven beyond ΔE 20), and that is published because a catalog stating its colours are approximate makes a different claim from one that does not. 1.4.0 (2026-08-18) — corrects 1.3.0, which said swatch images were read “for colorimetry only” while the shade rows shipped that day hotlink a per-variant product image. Two different acts, only one written down; both now stated. 1.3.0 (2026-08-18) — the five permitted brands’ scope widens from “no images” to per-variant swatch images read for colorimetry only, zero bytes persisted; cdn.shopify.com named as the host that serves them. Prices stay out of scope (superseded by 1.6.0). 1.2.0 (2026-08-17) — Section 2 split 4xx: a denied read (403) now fails closed like a 5xx, only a genuine 404 proceeds. Until this version Section 2 said a 4xx meant “no rules, crawl on” while Section 9 and brand-source-registry.json failed four 403 hosts closed; the document contradicted itself and the stricter half is now the written rule. 1.1.0 (2026-08-17) — Section 10 and four per-source rows added; 1.0.0 (2026-08-11) retained.
Owner
Backend catalog — UNASSIGNED.
Contact
crawler@looksmith.ai
Decision record
ADR-0009 (catalog ingestion). Referenced throughout; held internally.
Sibling
outbound-link-policy.md — what the app does with a link once the catalog holds one.

This document is the public statement of what the Looksmith catalog crawler does. Its governing rule: it claims LESS than the crawler does, and the crawler does EXACTLY what it claims. Reddit v. Anthropic (N.D. Cal. 2026) turned a public robots.txt claim contradicted by conduct into a Cal. Bus. & Prof. Code §17200 fraudulent-prong theory — so every sentence below is a commitment provable from the retained logs in Section 8, and no sentence is aspirational. No marketing claim about compliance is made here or anywhere else. Changes to this document bump the policy version; prior versions are retained.

  1. Identification
  2. robots.txt
  3. Rate policy
  4. What the crawler never does
  5. Images
  6. AI
  7. Takedown
  8. Evidence retention
  9. Per-source table
  10. Brand-direct URL resolution

Identification

Every crawler request carries the User-Agent string

LooksmithBot/1.0 (+https://looksmith.ai/crawler)

The string is never rotated, never blanked, never varied per host. This document is published verbatim at https://looksmith.ai/crawler before the first production batch runs (procurement gate, ADR-0009). Requests are signed under Web Bot Auth with Looksmith’s own bot-operator key pair — Shopify’s documented escalation path for identified bots, and the closest producible “authorization provided upon request” for the crawl tier.

robots.txt

Fetched before each batch, cached at most 24 hours, honoured literally. Each host’s robots.txt is fetched at the start of every batch and cached for the duration of that batch only — batch wall clock is ~60–90 minutes (ADR-0009), so no request is ever governed by a file older than 24 hours. The file is parsed per RFC 9309 for the token LooksmithBot, wildcards included. Crawl-delay is honoured literally even though RFC 9309 does not define it; where Crawl-delay exceeds the spacing in Section 3, the larger value wins.

Fetch failure fails CLOSED. A 5xx, timeout, TLS error, or unreachable host is treated as disallow-all for that host for the entire batch — the RFC 9309 §2.3.1.4 unreachable rule, applied as written. A robots fetch is never retried through its own failure within a batch.

A denied read is not an absent rule. Two different answers arrive as 4xx, and this policy separates them by what the status actually asserts:

Either way the status and the response body are retained in the batch snapshot (Section 8), because the difference between the two branches is the whole decision and it has to be auditable after the fact.

This is stricter than RFC 9309, deliberately, and the RFC is not claimed for it. §2.3.1.3 puts the entire 400–499 range in one bucket it calls “unavailable” and says that where the file is unavailable “the crawler MAY access any resources on the server” — to the RFC a 403 and a 404 are the same answer and both permit crawling. The MUST-assume-complete-disallow rule in §2.3.1.4 covers only server or network errors, the 500–599 range. So the 403 branch above is Looksmith’s rule, not the RFC’s: MAY is a permission this crawler declines to take when the reason the file is missing is that the host refused to serve it. Only the 5xx sentence above is the RFC applied as written.

This is also what the crawl set has actually done since 2026-08-17: the four 403 hosts in Section 9 and the unreadable verdict in db/seed/brand-source-registry.json are this rule, and resolve-brand-urls.ts enforces it in code (Section 10). Before 1.2.0 this paragraph said the opposite of both, which is exactly the claim-contradicted-by-conduct posture Section 1 exists to prevent — except inverted, with the document claiming a laxer rule than the code ran.

Every per-URL allow/deny decision is logged (Section 8). Shopify itself describes storefront robots.txt as “directional and advisory”; this crawler treats it as binding regardless.

Rate policy

What the crawler never does

Images

Merchant images are never re-hosted. A swatch image is fetched transiently into pipeline memory for colorimetry, and the derived value is a NUMBER — a LAB triplet — not pixels. Zero image bytes are written to durable storage; a test invariant enforces this (ADR-0009 Decision 3, 17 U.S.C. §106(1) intermediate-copy posture per Kelly v. Arriba Soft).

One offline pass departs from that, and says so (1.5.0). To colour the 569 shades whose brands publish no bare swatch, each shade’s own product photograph — the one already hotlinked in its catalog row — was fetched once and read by a vision model. A vision model needs a file, where every other pipeline decodes in memory, so those images were composed into contact sheets written to a scratch directory outside the repository, never committed, and removed after the run. Transient rather than zero, and durable only as one colour per shade. The invariant below is unaffected: it binds the catalog pipeline, and this is an offline curation tool that is not part of it.

The invariant is apps/api/test/catalog.zero-image-persistence.test.ts. It asserts that the crawl artifact contains no pixels, that every published image URL resolves to the merchant rather than to a Looksmith origin, and that the catalog module references no media origin, object store, or file write. Named here because a policy that claims an enforcement mechanism should say which one, and because this section said so for five days before the test existed.

AI

Crawled imagery is NEVER used to train, fine-tune, or improve any model, and NEVER enters the generation pipeline as a reference or conditioning input. The only model contact with a crawled image is transient inference-only colour reading whose sole output is the number in Section 5. Shopify API License §2.3.24 and Sephora Terms of Use §7 bind this commitment; it holds by construction, not by policy review.

Takedown

Notices go to crawler@looksmith.ai. On notice naming a source, crawling of that source STOPS within 24 hours (hard line (e), ADR-0009). Every source sits behind a per-source kill switch that removes it from the crawl set and from the published catalog without an app release; stop times and switch actuations are logged with timestamps. A robots.txt disallow appearing for LooksmithBot requires no notice: it takes effect at the next batch, no more than 7 days later.

Evidence retention

Each batch retains, per host: the fetched robots.txt bytes with timestamp and HTTP status; a snapshot of the host’s terms-of-use page (the source registry’s survey: assent mechanism, scraping clause, hotlink clause); every robots allow/deny decision; every 429/430 event and the cooldown action taken; and every kill-switch actuation. Nothing is rotated out: snapshots and decision logs are retained for the life of the catalog program.

Per-source table

Status as of 2026-08-11, with the robots observations dated 2026-08-17 in their own cells. Robots entries are live-fetch observations only — Shopify does not publish the default file’s contents, so no documentation claim is possible; entries marked UNVERIFIED are snapshotted from the first production batch onward (Section 8). A dated 403 here is a Section 2 fail-closed, not a note.

Host(s) What we fetch Headless robots.txt status Terms status (2026-08-11)
~320 brand DTC Shopify storefronts (registry enumerated in each batch artifact; incl. Credo, Bluemercury, Detox Market; e.g. glossier.com) GET /products.json; GET /products/{handle}.js for retained products only; PDP swatch markup No UNVERIFIED — per-batch snapshot Per-brand terms survey (assent, scraping clause, hotlink clause) REQUIRED before a brand enters the crawl set; Ajax API scope deviation recorded — the endpoint is documented for Shopify-hosted themes only (ADR-0009)
maccosmetics.com, clinique.com, bobbibrown.com, esteelauder.com, lamer.com PDP with the page’s own hydration executed, for HEX_VALUE_STRING (brand-authored hex). This list is the EXHAUSTIVE headless set Yes UNVERIFIED — per-batch snapshot Survey required before first batch
sephora.com via ac.cnstrc.com and chip URLs s{skuId}+sw.jpg Constructor.io search JSON (key ships unauthenticated in page source); 36×36 swatch chips (6/6 cold fetches 200, 2026-08-11) No www.sephora.com/robots.txt → 403 (AkamaiGHost Access Denied) to the Section 1 User-Agent, verified 2026-08-17. Fails closed under Section 2: www.sephora.com is not crawled. ac.cnstrc.com/robots.txt → 404 the same day, which is a genuine “no rules” and proceeds — but see the ruling note, because a third-party API host is not a way around an origin that refused us its rules Surveyed 2026-08-11: no assent clause (the Meta v. Bright Data posture); Terms of Use §7 bound per Section 6. OWNER RULING REQUIRED before the first production batch (2026-08-17): the 403 removes the “no readable prohibition” half of this row’s basis. The absent assent clause is unaffected — but under Section 2 the origin has declined to publish rules to us, so whether the ac.cnstrc.com route may still run is a decision, not an inference, and it is not made here
ulta.com NOT CRAWLED — — Ruling (owner, 2026-08-11): excluded — explicit assent plus a ban on “any collection and use of any product listings, descriptions, or prices”
5 brand storefronts surveyed permitted (rarebeauty.com, natashadenona.com, hourglasscosmetics.com, pixibeauty.com, lauramercier.com) GET /robots.txt; GET /products.json; GET /products/{handle}.js for candidate handles only. Out: URLs, GTINs, shade names, and the variant price with the instant it was observed. compare_at_price is not taken. Two image uses, both added 2026-08-18: the per-variant swatch is read for colorimetry (a colour value out, no bytes kept), and the per-variant product image URL is hotlinked — never fetched by us, never re-hosted. This cell read “no images, no prices” until 1.6.0, and by then both halves were false; the versions above record when each became so. No 200, honoured; unreadable fails closed Surveyed 2026-08-17 — no automated-access prohibition found. Registry: looksmith-backend/db/seed/brand-source-registry.json
tartecosmetics.com, maccosmetics.com, fentybeauty.com, stilacosmetics.com, patmcgrath.com, tomfordbeauty.com NOT CRAWLED — — Ruling (2026-08-17): excluded — explicit assent plus an automated-access ban, the same posture that excluded ulta.com. tarte’s reaches “any robot, spider, or other automatic device… for any purpose”
anastasiabeverlyhills.com NOT CRAWLED — — Ruling (2026-08-17): excluded — no access ban, but reuse of site content “for public or commercial purposes without written permission” is barred, which is what a commercial catalog does. Also bars downloading product images outright
benefitcosmetics.com, charlottetilbury.com, bobbibrowncosmetics.com, makeupforever.com NOT CRAWLED — 403 to any non-browser client (all four re-verified 2026-08-17) Failed closed. Absence of a readable policy is not consent
Tokenless Storefront GraphQL (any host) NOT USED — — Shopify API License §2.3.14 bans systematic automated collection; Shopify is the counterparty and can block at platform level

Brand-direct URL resolution (2026-08-17)

The retailer crawl yields a retailer slug per row. A brand’s own product URL is not derivable from it, so linking a shopper to the brand rather than the retailer requires resolving each row against the brand’s storefront — the Shopify row of the table above, whose stated precondition is a per-brand terms survey.

What the survey found. Of twelve brands considered, six ban automated access outright, one bars commercial reuse of site content, and four are unreachable to any non-browser client. Five are permitted. The verdicts, with the prohibiting sentence quoted verbatim for each exclusion, are in looksmith-backend/db/seed/brand-source-registry.json.

The registry is the gate, not a record of one: scripts/resolve-brand-urls.ts reads it and crawls only permitted brands, a brand absent from it is not crawled either, and catalog.brand-source-registry.test.ts asserts both that the resolver has no hardcoded host list and that the resolved cache contains no non-permitted host. A survey that lives only in prose is the failure mode Section 1 of this policy exists to avoid.

The resolver’s robots handling is Section 2’s 403 branch, and is in fact stricter: permitsResolution() treats every non-2xx robots response as a refusal, so it skips a 404 host that Section 2 would let the batch crawler proceed against. Recorded rather than relaxed — a resolver that reads twelve storefronts by hand-audited GTIN has no need of the 404 permission, and the gap runs in the safe direction.

Resolution is by GTIN only. 97% of crawled rows carry a GTIN; a brand’s products/{handle}.js carries the barcode and the variant id, so a match is an identity match and the resulting URL pins the exact shade. Name matching is refused outright: on this data it sent “Velvet Teddy” to a blush and “Diva” to “Big Diva Energy” blush, because shade names collide across product lines. A Buy button that opens the wrong product is worse than one that opens a stale page.

Coverage, stated plainly. 120 of 2,011 rows — 6% — resolve to a brand URL under this policy. The remaining 94% is not a technical gap: it is brands whose terms forbid this, and the licensed route (an affiliate or distributor product feed, which carries the merchant’s own URL and the GTIN together) is the way to reach them. Crawling harder is not.

Conduct gap, recorded. maccosmetics.com/products.json was read ~10 times on 2026-08-17, and the resolver was run against MAC and Rare Beauty, BEFORE this survey existed. MAC prohibits automated access; Rare Beauty does not. The MAC reads were used to repair nine catalog buy URLs that a platform migration had broken, one of which was already a 404. Those nine URLs are ordinary public facts a person could have looked up by hand and they remain in the seed, but the method preceded its authorisation and the ordering was wrong. MAC is excluded from automated resolution from this date. This paragraph exists because a policy whose breaches are unrecorded is the Reddit failure mode described in Section 1.