[{"data":1,"prerenderedAt":1726},["ShallowReactive",2],{"api-blog-best-brand-data-apis-2026":3,"api-blog-related-best-brand-data-apis-2026":448},{"id":4,"title":5,"author":6,"body":7,"category":416,"date":417,"dateModified":418,"description":419,"extension":420,"faqs":421,"image":418,"meta":432,"navigation":433,"path":434,"readingTime":435,"relatedSlugs":436,"seo":440,"stem":441,"tags":442,"__hash__":447},"apiBlog/api-blog/best-brand-data-apis-2026.md","Best Brand Data APIs in 2026 (Logos, Colors, Fonts)","Soraia",{"type":8,"value":9,"toc":401},"minimark",[10,19,22,27,37,45,67,71,79,96,100,107,124,128,146,196,199,221,225,342,346,359,363,368,371,375,378,382,385,389,397],[11,12,13,14,18],"p",{},"If you need to turn ",[15,16,17],"code",{},"stripe.com"," into a logo, brand colors, and fonts inside your product, four options are worth your time in 2026. The short version: logos alone are a solved, free problem; the money question is how much design depth you need behind the logo.",[11,20,21],{},"The reason this category exists is not design tooling, it is conversion. Brandfetch's customer pages cite Typeform at +5% free-to-paid and Senja at +233% activation after wiring brand personalization into onboarding. \"Enter your URL, watch the product become yours\" is one of the few onboarding tricks with published numbers behind it.",[23,24,26],"h2",{"id":25},"_1-brandfetch-the-category-owner","1. Brandfetch - the category owner",[11,28,29,36],{},[30,31,35],"a",{"href":32,"rel":33},"https://brandfetch.com",[34],"nofollow","Brandfetch"," is the default answer, with Meta, Typeform, and Lucid on its logo wall. One call returns logos in several formats, brand colors, fonts, and company info.",[11,38,39,40,44],{},"Their pricing (verified July 2026): free tier of 100 Brand API requests, then the Brand API at ",[41,42,43],"strong",{},"$99/mo for 2,500 calls"," with $0.10 overage. The Logo API alone is free up to 500k requests a month, which is quietly the best deal in the category.",[46,47,48,55,61],"ul",{},[49,50,51,54],"li",{},[41,52,53],{},"Good:"," the deepest logo coverage in the market, hand-curated for big brands, reliable.",[49,56,57,60],{},[41,58,59],{},"Bad:"," $99/mo is real money for 2,500 calls, and the design data stops at \"a few colors and font names\". You cannot theme a UI from it - there is no spacing, no type scale, no shadows.",[49,62,63,66],{},[41,64,65],{},"Use it when:"," logo quality and coverage matter more than design depth, and the budget is there.",[23,68,70],{"id":69},"_2-logodev-the-pragmatic-logo-endpoint","2. Logo.dev - the pragmatic logo endpoint",[11,72,73,78],{},[30,74,77],{"href":75,"rel":76},"https://logo.dev",[34],"Logo.dev"," does one thing: domain in, logo image out, with a generous free tier and simple img-tag integration.",[46,80,81,86,91],{},[49,82,83,85],{},[41,84,53],{}," dead simple, fast to integrate, free for most usage.",[49,87,88,90],{},[41,89,59],{}," logos only. No colors, no fonts, no JSON brand object.",[49,92,93,95],{},[41,94,65],{}," you literally just need company logos in a table or CRM.",[23,97,99],{"id":98},"_3-clearbit-logo-api-the-old-free-standby","3. Clearbit Logo API - the old free standby",[11,101,102,103,106],{},"Clearbit's ",[15,104,105],{},"logo.clearbit.com/{domain}"," endpoint has been free and everywhere for a decade. Clearbit is now part of HubSpot, and the endpoint has outlived several announcements about its future.",[46,108,109,114,119],{},[49,110,111,113],{},[41,112,53],{}," free, zero setup, still works.",[49,115,116,118],{},[41,117,59],{}," logos only, no SLA, and it lives at an acquirer's pleasure. I would not build a paid product on it in 2026.",[49,120,121,123],{},[41,122,65],{}," internal tools and prototypes.",[23,125,127],{"id":126},"_4-miromiro-brand-extract-api-identity-plus-the-design-system","4. MiroMiro brand + extract API - identity plus the design system",[11,129,130,131,135,136,140,141,145],{},"Our ",[30,132,134],{"href":133},"/api/docs/brand","/v1/brand"," endpoint returns the logo, palette, and fonts like a brand API. The difference is what sits next to it: ",[30,137,139],{"href":138},"/api/docs/design-tokens","/v1/extract"," returns the site's working design system - ranked colors, type scale, spacing, radii, shadows, gradients, motion - and ",[30,142,144],{"href":143},"/api/docs/code","/v1/code"," turns a section into Tailwind/JSX/Vue. Same key, same call shape:",[147,148,153],"pre",{"className":149,"code":150,"language":151,"meta":152,"style":152},"language-bash shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","curl \"https://miromiro.app/api/v1/brand?url=stripe.com&fields=logo,palette\" \\\n  -H \"Authorization: Bearer $MIROMIRO_API_KEY\"\n","bash","",[15,154,155,179],{"__ignoreMap":152},[156,157,160,164,168,172,175],"span",{"class":158,"line":159},"line",1,[156,161,163],{"class":162},"sBMFI","curl",[156,165,167],{"class":166},"sMK4o"," \"",[156,169,171],{"class":170},"sfazB","https://miromiro.app/api/v1/brand?url=stripe.com&fields=logo,palette",[156,173,174],{"class":166},"\"",[156,176,178],{"class":177},"sTEyZ"," \\\n",[156,180,182,185,187,190,193],{"class":158,"line":181},2,[156,183,184],{"class":170},"  -H",[156,186,167],{"class":166},[156,188,189],{"class":170},"Authorization: Bearer ",[156,191,192],{"class":177},"$MIROMIRO_API_KEY",[156,194,195],{"class":166},"\"\n",[11,197,198],{},"Pricing: 100 free credits a month with no card, then from €19/mo for 5,000 credits. A brand call costs 15 credits, a token extraction 10.",[46,200,201,211,216],{},[49,202,203,205,206,210],{},[41,204,53],{}," the only option here that returns enough to actually ",[207,208,209],"em",{},"theme"," a product, not just badge it. Built for the \"enter your URL\" onboarding flow.",[49,212,213,215],{},[41,214,59],{}," our logo curation is not Brandfetch's - we extract from the live site rather than maintain a hand-checked logo library, so obscure-brand edge cases favor them. Static HTML + CSS parsing; no JS execution.",[49,217,218,220],{},[41,219,65],{}," the goal is brand-personalized UI (onboarding, generated apps, email builders), not a logo lookup.",[23,222,224],{"id":223},"side-by-side","Side by side",[226,227,228,256],"table",{},[229,230,231],"thead",{},[232,233,234,238,241,244,247,250,253],"tr",{},[235,236,237],"th",{},"API",[235,239,240],{},"Logo",[235,242,243],{},"Colors",[235,245,246],{},"Fonts",[235,248,249],{},"Full design system",[235,251,252],{},"Free tier",[235,254,255],{},"Paid entry",[257,258,259,282,301,320],"tbody",{},[232,260,261,264,267,270,273,276,279],{},[262,263,35],"td",{},[262,265,266],{},"Excellent",[262,268,269],{},"Basic",[262,271,272],{},"Names",[262,274,275],{},"No",[262,277,278],{},"100 req",[262,280,281],{},"$99/mo (2,500 calls)",[232,283,284,286,289,291,293,295,298],{},[262,285,77],{},[262,287,288],{},"Good",[262,290,275],{},[262,292,275],{},[262,294,275],{},[262,296,297],{},"Generous",[262,299,300],{},"Cheap",[232,302,303,306,308,310,312,314,317],{},[262,304,305],{},"Clearbit Logo",[262,307,288],{},[262,309,275],{},[262,311,275],{},[262,313,275],{},[262,315,316],{},"Unlimited-ish",[262,318,319],{},"n/a",[232,321,322,325,327,330,333,336,339],{},[262,323,324],{},"MiroMiro",[262,326,288],{},[262,328,329],{},"Ranked, full palette",[262,331,332],{},"Families + weights + sizes",[262,334,335],{},"Yes (tokens + section code)",[262,337,338],{},"100 credits/mo",[262,340,341],{},"€19/mo (5,000 credits)",[23,343,345],{"id":344},"the-verdict","The verdict",[11,347,348,349,353,354,358],{},"Logos only: Logo.dev, or Clearbit if it is internal. Maximum logo coverage with budget: Brandfetch. Theming a product to the user's brand: that is the job we built for - a logo plus three colors will not carry an onboarding flow that promises \"your brand, applied\". You can ",[30,350,352],{"href":351},"/api/dashboard/keys","get a free key"," and test your own domain in the ",[30,355,357],{"href":356},"/api/demo","playground"," in under a minute.",[23,360,362],{"id":361},"frequently-asked-questions","Frequently asked questions",[364,365,367],"h3",{"id":366},"what-is-a-brand-data-api","What is a brand data API?",[11,369,370],{},"Domain in, structured brand assets out: logo, colors, fonts. Products use them for onboarding personalization, CRM enrichment, and auto-theming.",[364,372,374],{"id":373},"is-there-a-free-logo-api","Is there a free logo API?",[11,376,377],{},"Yes, two: Brandfetch's Logo API tier (free to 500k req/mo) and Clearbit's endpoint. For logos alone, pay nothing.",[364,379,381],{"id":380},"why-fetch-brand-data-during-onboarding","Why fetch brand data during onboarding?",[11,383,384],{},"Published numbers: +5% free-to-paid at Typeform, +233% activation at Senja (per Brandfetch's customer pages). Users activate faster inside a product that already looks like theirs.",[364,386,388],{"id":387},"brand-api-vs-design-token-api","Brand API vs design token API?",[11,390,391,392,396],{},"Identity vs system. A logo and five hex codes badge a UI; a token set (palette, type scale, spacing, shadows, motion) is what you need to build one. The ",[30,393,395],{"href":394},"/api/blog/design-token-types-reference","design token types reference"," breaks down the difference.",[398,399,400],"style",{},"html pre.shiki code .sBMFI, html code.shiki .sBMFI{--shiki-light:#E2931D;--shiki-default:#FFCB6B;--shiki-dark:#FFCB6B}html pre.shiki code .sMK4o, html code.shiki .sMK4o{--shiki-light:#39ADB5;--shiki-default:#89DDFF;--shiki-dark:#89DDFF}html pre.shiki code .sfazB, html code.shiki .sfazB{--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D}html pre.shiki code .sTEyZ, html code.shiki .sTEyZ{--shiki-light:#90A4AE;--shiki-default:#EEFFFF;--shiki-dark:#BABED8}html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":152,"searchDepth":181,"depth":181,"links":402},[403,404,405,406,407,408,409],{"id":25,"depth":181,"text":26},{"id":69,"depth":181,"text":70},{"id":98,"depth":181,"text":99},{"id":126,"depth":181,"text":127},{"id":223,"depth":181,"text":224},{"id":344,"depth":181,"text":345},{"id":361,"depth":181,"text":362,"children":410},[411,413,414,415],{"id":366,"depth":412,"text":367},3,{"id":373,"depth":412,"text":374},{"id":380,"depth":412,"text":381},{"id":387,"depth":412,"text":388},"Comparisons","2026-07-18",null,"Brandfetch, Logo.dev, Clearbit, and MiroMiro compared honestly on depth and pricing - when a logo API is enough and when you need the full design system.","md",[422,424,426,429],{"question":367,"answer":423},"An API that takes a domain name and returns the company's brand assets - typically the logo, brand colors, and fonts - as structured data. Products use them to personalize onboarding, enrich CRM records, and theme UIs to a customer's brand automatically.",{"question":374,"answer":425},"Yes. Brandfetch's Logo API tier is free up to 500,000 requests a month, and Clearbit's long-standing logo endpoint (now part of HubSpot) has been free for years. For logos alone, you should not be paying.",{"question":427,"answer":428},"Why do products fetch brand data during onboarding?","Because personalization measurably converts. Brandfetch's own customer pages cite Typeform seeing +5% free-to-paid and Senja +233% activation after adding brand personalization to onboarding. A user who sees their own brand in your product is closer to activated.",{"question":430,"answer":431},"What is the difference between a brand API and a design token API?","A brand API returns identity: logo, a handful of colors, maybe fonts. A design token API returns the working design system: the full palette with usage ranking, type scale, spacing, radii, shadows, and motion. If you are theming a UI, identity alone is not enough to build with.",{},true,"/api-blog/best-brand-data-apis-2026",7,[437,438,439],"best-design-token-extraction-tools-2026","best-firecrawl-alternatives-2026","how-to-extract-brand-assets-in-bulk",{"title":5,"description":419},"api-blog/best-brand-data-apis-2026",[443,444,445,446],"brand api","brandfetch","logo api","comparison","6oZS0krURBt4VY9_ej0hmKJd6CMEe4ynuxFS-SMRqIw",[449,884,1240],{"id":450,"title":451,"author":6,"body":452,"category":416,"date":417,"dateModified":418,"description":858,"extension":420,"faqs":859,"image":418,"meta":871,"navigation":433,"path":872,"readingTime":435,"relatedSlugs":873,"seo":877,"stem":878,"tags":879,"__hash__":883},"apiBlog/api-blog/best-design-token-extraction-tools-2026.md","Best Design Token Extraction Tools in 2026",{"type":8,"value":453,"toc":842},[454,457,461,469,486,490,498,515,519,527,544,548,560,577,581,588,605,609,612,642,663,667,781,783,793,795,799,802,806,820,824,831,835,840],[11,455,456],{},"There are six honest ways to get design tokens out of a live website in 2026, and the right one depends on whether you are doing this once, weekly, or inside a product. Here is the list, with the tradeoffs stated plainly. Yes, our own API is on it - ranked where I would rank it if it were not ours.",[23,458,460],{"id":459},"_1-css-stats-fastest-free-look","1. CSS Stats - fastest free look",[11,462,463,468],{},[30,464,467],{"href":465,"rel":466},"https://cssstats.com",[34],"cssstats.com"," has been the \"paste a URL, see the CSS\" tool for a decade. It reports every color, font size, and font family the stylesheet declares, plus specificity graphs nobody asks for but everyone screenshots.",[46,470,471,476,481],{},[49,472,473,475],{},[41,474,53],{}," free, no signup, instant.",[49,477,478,480],{},[41,479,59],{}," it reads declared CSS, not used CSS. A site shipping a big framework stylesheet reports hundreds of colors that never render. There is no export worth automating against.",[49,482,483,485],{},[41,484,65],{}," you want a two-minute overview of a site's CSS health, not a token set.",[23,487,489],{"id":488},"_2-project-wallace-the-analyzer-that-kept-evolving","2. Project Wallace - the analyzer that kept evolving",[11,491,492,497],{},[30,493,496],{"href":494,"rel":495},"https://www.projectwallace.com",[34],"Project Wallace"," does what CSS Stats does but deeper: color counts, typography breakdowns, and change tracking over time if you create an account.",[46,499,500,505,510],{},[49,501,502,504],{},[41,503,53],{}," the most thorough free CSS analysis on the web, actively maintained.",[49,506,507,509],{},[41,508,59],{}," same fundamental limit - declared, not rendered. Still a report you read, not a token file you use.",[49,511,512,514],{},[41,513,65],{}," auditing your own CSS over time, or studying how a site's stylesheet is put together.",[23,516,518],{"id":517},"_3-superposition-tokens-on-your-desktop","3. Superposition - tokens on your desktop",[11,520,521,526],{},[30,522,525],{"href":523,"rel":524},"https://superposition.design",[34],"Superposition"," is a desktop app that crawls a page and exports design tokens to CSS variables, Sass, JS, and even Adobe XD. It was ahead of its time.",[46,528,529,534,539],{},[49,530,531,533],{},[41,532,53],{}," genuine token export formats, one-click, offline app.",[49,535,536,538],{},[41,537,59],{}," desktop-only, so there is nothing to call from code or CI. Development has been quiet for a while - check that the export you need still works.",[49,540,541,543],{},[41,542,65],{}," you are a designer who wants tokens from a reference site without touching a terminal.",[23,545,547],{"id":546},"_4-a-diy-playwright-script-full-control","4. A DIY Playwright script - full control",[11,549,550,551,554,555,559],{},"Load the page in a headless browser, read ",[15,552,553],{},"getComputedStyle"," off every element, count what repeats. We wrote the complete walkthrough in ",[30,556,558],{"href":557},"/api/blog/how-to-extract-design-tokens-from-any-website-in-python","How to Extract Design Tokens from Any Website in Python",".",[46,561,562,567,572],{},[49,563,564,566],{},[41,565,53],{}," free, sees exactly what renders (including runtime CSS-in-JS), infinitely customizable.",[49,568,569,571],{},[41,570,59],{}," you now run browser infrastructure. Cookie banners, lazy fonts, and viewport-dependent styles are your problem, and the script that worked in April breaks in July.",[49,573,574,576],{},[41,575,65],{}," one-off extractions, research, or when you need something no tool exposes.",[23,578,580],{"id":579},"_5-miromiro-browser-extension-point-and-click","5. MiroMiro browser extension - point and click",[11,582,130,583,587],{},[30,584,586],{"href":585},"/features/design-tokens","extension"," shows a site's tokens live as you browse: colors, typography, spacing, assets. Click an element, see its resolved styles, export what you want.",[46,589,590,595,600],{},[49,591,592,594],{},[41,593,53],{}," zero setup, visual, great for picking tokens from a specific component rather than the whole page.",[49,596,597,599],{},[41,598,59],{}," it is a human tool. Nothing to automate, no JSON endpoint, and it extracts what you look at, not a thousand URLs.",[49,601,602,604],{},[41,603,65],{}," hand-picking design decisions from reference sites during design or build.",[23,606,608],{"id":607},"_6-miromiro-extract-api-tokens-as-an-endpoint","6. MiroMiro extract API - tokens as an endpoint",[11,610,611],{},"One GET request returns the ranked token set as JSON, CSS variables, or a Tailwind theme block:",[147,613,615],{"className":149,"code":614,"language":151,"meta":152,"style":152},"curl \"https://miromiro.app/api/v1/extract?url=stripe.com&format=tailwind\" \\\n  -H \"Authorization: Bearer $MIROMIRO_API_KEY\"\n",[15,616,617,630],{"__ignoreMap":152},[156,618,619,621,623,626,628],{"class":158,"line":159},[156,620,163],{"class":162},[156,622,167],{"class":166},[156,624,625],{"class":170},"https://miromiro.app/api/v1/extract?url=stripe.com&format=tailwind",[156,627,174],{"class":166},[156,629,178],{"class":177},[156,631,632,634,636,638,640],{"class":158,"line":181},[156,633,184],{"class":170},[156,635,167],{"class":166},[156,637,189],{"class":170},[156,639,192],{"class":177},[156,641,195],{"class":166},[46,643,644,653,658],{},[49,645,646,648,649,652],{},[41,647,53],{}," deduped and ranked (not raw declared values), covers colors, type scale, spacing, radii, shadows, gradients, and motion, resolves ",[15,650,651],{},"var()"," chains and every declared breakpoint. It is the only option on this list you can put inside a product.",[49,654,655,657],{},[41,656,59],{}," it parses static HTML + CSS and does not execute JavaScript, so a fully client-rendered app can come back thin. Costs credits past the free 100 a month.",[49,659,660,662],{},[41,661,65],{}," tokens need to arrive programmatically - agent pipelines, onboarding flows, bulk analysis.",[23,664,666],{"id":665},"the-comparison-in-one-table","The comparison in one table",[226,668,669,688],{},[229,670,671],{},[232,672,673,676,679,682,685],{},[235,674,675],{},"Tool",[235,677,678],{},"Output",[235,680,681],{},"Rendered or declared?",[235,683,684],{},"Automatable",[235,686,687],{},"Price",[257,689,690,706,720,734,750,764],{},[232,691,692,695,698,701,703],{},[262,693,694],{},"CSS Stats",[262,696,697],{},"Web report",[262,699,700],{},"Declared",[262,702,275],{},[262,704,705],{},"Free",[232,707,708,710,713,715,718],{},[262,709,496],{},[262,711,712],{},"Web report + tracking",[262,714,700],{},[262,716,717],{},"Partially",[262,719,252],{},[232,721,722,724,727,730,732],{},[262,723,525],{},[262,725,726],{},"Token files (CSS/Sass/JS)",[262,728,729],{},"Rendered",[262,731,275],{},[262,733,705],{},[232,735,736,739,742,744,747],{},[262,737,738],{},"DIY Playwright",[262,740,741],{},"Whatever you build",[262,743,729],{},[262,745,746],{},"Yes, you maintain it",[262,748,749],{},"Your time",[232,751,752,755,758,760,762],{},[262,753,754],{},"MiroMiro extension",[262,756,757],{},"Visual + export",[262,759,729],{},[262,761,275],{},[262,763,252],{},[232,765,766,769,772,775,778],{},[262,767,768],{},"MiroMiro API",[262,770,771],{},"JSON / CSS vars / Tailwind",[262,773,774],{},"Parsed + resolved",[262,776,777],{},"Yes",[262,779,780],{},"100 credits/mo free, then from €19/mo",[23,782,345],{"id":344},[11,784,785,786,789,790,792],{},"For a quick look, Project Wallace. For hand-picking from a reference site, the extension. For anything that runs on a schedule or inside a product, the API - and I would say that even if it were someone else's, because it is the only entry here that returns a token file over HTTP. You can ",[30,787,788],{"href":351},"get a free API key"," (100 credits a month, no card) or poke it in the ",[30,791,357],{"href":356}," first.",[23,794,362],{"id":361},[364,796,798],{"id":797},"what-is-the-fastest-free-way-to-extract-design-tokens-from-a-website","What is the fastest free way to extract design tokens from a website?",[11,800,801],{},"Paste the URL into CSS Stats or Project Wallace. Seconds, no signup. Expect raw declared values, not a cleaned token set.",[364,803,805],{"id":804},"can-i-get-tokens-as-css-variables-or-a-tailwind-config-directly","Can I get tokens as CSS variables or a Tailwind config directly?",[11,807,808,809,812,813,816,817,559],{},"Yes - ",[15,810,811],{},"format=css"," returns custom properties, ",[15,814,815],{},"format=tailwind"," returns a theme block. Details in the ",[30,818,819],{"href":138},"design tokens docs",[364,821,823],{"id":822},"do-these-tools-work-on-tailwind-sites","Do these tools work on Tailwind sites?",[11,825,826,827,830],{},"The ones reading real CSS or computed styles do; ",[15,828,829],{},"bg-indigo-600"," is just CSS once shipped. Tools that only look for CSS custom properties miss most of a Tailwind site's system.",[364,832,834],{"id":833},"what-is-the-difference-between-a-token-set-and-a-color-palette","What is the difference between a token set and a color palette?",[11,836,837,838,559],{},"A palette is one token group. A full token set adds the type scale, spacing, radii, shadows, motion, and z-index layers - see the ",[30,839,395],{"href":394},[398,841,400],{},{"title":152,"searchDepth":181,"depth":181,"links":843},[844,845,846,847,848,849,850,851,852],{"id":459,"depth":181,"text":460},{"id":488,"depth":181,"text":489},{"id":517,"depth":181,"text":518},{"id":546,"depth":181,"text":547},{"id":579,"depth":181,"text":580},{"id":607,"depth":181,"text":608},{"id":665,"depth":181,"text":666},{"id":344,"depth":181,"text":345},{"id":361,"depth":181,"text":362,"children":853},[854,855,856,857],{"id":797,"depth":412,"text":798},{"id":804,"depth":412,"text":805},{"id":822,"depth":412,"text":823},{"id":833,"depth":412,"text":834},"Six real ways to pull design tokens from a live website in 2026 - free analyzers, a DIY script, a browser extension, and an API - honestly ranked.",[860,862,865,868],{"question":798,"answer":861},"Paste the URL into a CSS analyzer like CSS Stats or Project Wallace. You get colors, font sizes, and specificity data in seconds with no signup. The output is raw declared values rather than a cleaned token set, so expect to filter it by hand.",{"question":863,"answer":864},"Can I get design tokens as CSS variables or a Tailwind config directly?","Yes. The MiroMiro extract endpoint takes format=css to return CSS custom properties or format=tailwind to return a theme config block, so the tokens drop straight into a codebase without a conversion step.",{"question":866,"answer":867},"Do design token extractors work on Tailwind sites?","Tools that read the site's actual CSS or computed styles do, because bg-indigo-600 is just CSS once shipped. Tools that look for CSS variables only will miss most of a Tailwind site's system, since utility-first sites often define few custom properties.",{"question":869,"answer":870},"What is the difference between extracting design tokens and extracting a color palette?","A palette is one token group. A design token set also covers the type scale, spacing, radii, shadows, motion, and z-index layers - the full set of decisions you need to rebuild UI in the site's style.",{},"/api-blog/best-design-token-extraction-tools-2026",[874,875,876],"how-to-extract-design-tokens-from-any-website-in-python","design-token-types-reference","best-design-to-code-apis-2026",{"title":451,"description":858},"api-blog/best-design-token-extraction-tools-2026",[880,881,446,882],"design tokens","tools","css analysis","GhGcQEA8Q7VmKm_mMFRSpL45k0PaBwatHPzBRiFUK9g",{"id":885,"title":886,"author":6,"body":887,"category":416,"date":417,"dateModified":418,"description":1215,"extension":420,"faqs":1216,"image":418,"meta":1228,"navigation":433,"path":1229,"readingTime":435,"relatedSlugs":1230,"seo":1232,"stem":1233,"tags":1234,"__hash__":1239},"apiBlog/api-blog/best-firecrawl-alternatives-2026.md","5 Best Firecrawl Alternatives in 2026",{"type":8,"value":888,"toc":1200},[889,892,895,899,907,925,929,941,958,962,974,991,995,1003,1020,1024,1031,1061,1067,1084,1088,1155,1157,1168,1170,1174,1177,1181,1184,1188,1191,1195,1198],[11,890,891],{},"\"Firecrawl alternative\" usually means one of five different complaints: it costs too much at volume, you want open source, you need heavier browser automation, you want search instead of crawling, or you need something it does not extract at all. Different complaint, different answer - here are the five, honestly.",[11,893,894],{},"For the record, Firecrawl is good. Their pricing (verified July 2026): free 1,000 credits, Hobby $16/mo for 5,000 pages, Standard $83/mo for 100k, up to Scale at $599/mo for a million. Simple unit, fair prices. If it is doing its job for you, keep it, and skip to the last section for the one case where it genuinely cannot help.",[23,896,898],{"id":897},"_1-crawl4ai-the-open-source-answer","1. Crawl4AI - the open-source answer",[11,900,901,906],{},[30,902,905],{"href":903,"rel":904},"https://github.com/unclecode/crawl4ai",[34],"Crawl4AI"," is the most popular OSS route: a Python crawler built for LLM pipelines, outputting clean markdown, with an enthusiastic community and rapid development.",[46,908,909,914,919],{},[49,910,911,913],{},[41,912,53],{}," free forever, self-hosted, no per-page anxiety, full data control.",[49,915,916,918],{},[41,917,59],{}," you own the infrastructure - proxies, JS rendering, anti-bot upkeep, retries. The afternoon it saves in fees it can take back in maintenance.",[49,920,921,924],{},[41,922,923],{},"Pick it when:"," volume is high, budgets are tight, and you have the engineering appetite.",[23,926,928],{"id":927},"_2-jina-reader-the-simplest-possible-api","2. Jina Reader - the simplest possible API",[11,930,931,936,937,940],{},[30,932,935],{"href":933,"rel":934},"https://jina.ai/reader/",[34],"Jina Reader"," is the lowest-friction markdown-from-URL service that exists: prefix a URL with ",[15,938,939],{},"r.jina.ai/"," and get LLM-ready text, generous free tier, no signup for casual use.",[46,942,943,948,953],{},[49,944,945,947],{},[41,946,53],{}," zero setup, genuinely free for light use, great output for articles and docs.",[49,949,950,952],{},[41,951,59],{}," single-page reads, not site crawling; fewer knobs when a page fights back.",[49,954,955,957],{},[41,956,923],{}," your agent reads pages one at a time and you want the cheapest possible path.",[23,959,961],{"id":960},"_3-apify-when-scraping-is-an-application","3. Apify - when scraping is an application",[11,963,964,969,970,559],{},[30,965,968],{"href":966,"rel":967},"https://apify.com",[34],"Apify"," is a platform, not an endpoint: thousands of prebuilt scrapers (\"Actors\") for specific sites, scheduling, storage, and a marketplace. We compared it in depth in our ",[30,971,973],{"href":972},"/api/compare/apify-alternative","Apify alternative page",[46,975,976,981,986],{},[49,977,978,980],{},[41,979,53],{}," the prebuilt Actors cover sites that generic crawlers struggle with; real orchestration features.",[49,982,983,985],{},[41,984,59],{}," platform complexity and platform pricing. For \"give me this page as markdown\" it is a lot of machinery.",[49,987,988,990],{},[41,989,923],{}," scraping specific difficult sites on a schedule is a core workflow.",[23,992,994],{"id":993},"_4-tavily-search-instead-of-crawl","4. Tavily - search instead of crawl",[11,996,997,1002],{},[30,998,1001],{"href":999,"rel":1000},"https://tavily.com",[34],"Tavily"," reframes the problem: instead of crawling sites, it gives agents a search API that returns answer-shaped, cited results.",[46,1004,1005,1010,1015],{},[49,1006,1007,1009],{},[41,1008,53],{}," for research agents, search-shaped input beats crawl-shaped input - less context, more signal.",[49,1011,1012,1014],{},[41,1013,59],{}," it is not extraction. You cannot point it at a specific page and get its full content faithfully.",[49,1016,1017,1019],{},[41,1018,923],{}," the agent's question is \"what is true about X\" rather than \"what does this page say\".",[23,1021,1023],{"id":1022},"_5-miromiro-when-the-answer-isnt-text-at-all","5. MiroMiro - when the answer isn't text at all",[11,1025,1026,1027,1030],{},"Everything above ships text. If the reason you are searching for alternatives is that your agent builds ",[41,1028,1029],{},"UI"," - an app generator theming output to a user's brand, an onboarding flow that imports a company's look, a coding agent that needs real design context - then markdown was never the right output format, and that is the lane we occupy:",[147,1032,1034],{"className":149,"code":1033,"language":151,"meta":152,"style":152},"curl \"https://miromiro.app/api/v1/extract?url=stripe.com\" \\\n  -H \"Authorization: Bearer $MIROMIRO_API_KEY\"\n",[15,1035,1036,1049],{"__ignoreMap":152},[156,1037,1038,1040,1042,1045,1047],{"class":158,"line":159},[156,1039,163],{"class":162},[156,1041,167],{"class":166},[156,1043,1044],{"class":170},"https://miromiro.app/api/v1/extract?url=stripe.com",[156,1046,174],{"class":166},[156,1048,178],{"class":177},[156,1050,1051,1053,1055,1057,1059],{"class":158,"line":181},[156,1052,184],{"class":170},[156,1054,167],{"class":166},[156,1056,189],{"class":170},[156,1058,192],{"class":177},[156,1060,195],{"class":166},[11,1062,1063,1064,1066],{},"Design tokens (ranked colors, type scale, spacing, shadows, motion), brand data, SVGs, fonts, and section-level Tailwind/JSX/Vue via ",[30,1065,144],{"href":143},". MCP server included for Claude Code and Cursor. Free tier is 100 credits a month, no card; paid from €19/mo.",[46,1068,1069,1074,1079],{},[49,1070,1071,1073],{},[41,1072,53],{}," the only entry here that returns a design system. Complements Firecrawl rather than replacing it - content from them, design from us is a real pattern.",[49,1075,1076,1078],{},[41,1077,59],{}," it will not crawl a thousand documentation pages into markdown. Wrong tool for that; use the four above.",[49,1080,1081,1083],{},[41,1082,923],{}," the deliverable is on-brand UI, not text.",[23,1085,1087],{"id":1086},"the-five-sorted-by-complaint","The five, sorted by complaint",[226,1089,1090,1103],{},[229,1091,1092],{},[232,1093,1094,1097,1100],{},[235,1095,1096],{},"Your complaint with Firecrawl",[235,1098,1099],{},"Alternative",[235,1101,1102],{},"Why",[257,1104,1105,1115,1125,1135,1145],{},[232,1106,1107,1110,1112],{},[262,1108,1109],{},"Cost at volume",[262,1111,905],{},[262,1113,1114],{},"Self-hosted, no per-page fee",[232,1116,1117,1120,1122],{},[262,1118,1119],{},"Just want one page as markdown",[262,1121,935],{},[262,1123,1124],{},"URL prefix, done",[232,1126,1127,1130,1132],{},[262,1128,1129],{},"Difficult specific sites, scheduling",[262,1131,968],{},[262,1133,1134],{},"Prebuilt Actors + orchestration",[232,1136,1137,1140,1142],{},[262,1138,1139],{},"Agent needs answers, not pages",[262,1141,1001],{},[262,1143,1144],{},"Search-shaped API",[232,1146,1147,1150,1152],{},[262,1148,1149],{},"Need design, not text",[262,1151,324],{},[262,1153,1154],{},"Tokens, brand, assets, section code",[23,1156,345],{"id":344},[11,1158,1159,1160,1163,1164,559],{},"For text at scale, Crawl4AI or stay with Firecrawl - their pricing is honestly hard to beat below 100k pages. For agent research, Tavily. And if you landed here because no amount of markdown gets you a brand's colors, fonts, and components: that is not a Firecrawl weakness, it is a category boundary, and it is the one we built for. ",[30,1161,1162],{"href":351},"Free key",", or the full head-to-head is at ",[30,1165,1167],{"href":1166},"/api/compare/firecrawl-alternative","MiroMiro vs Firecrawl",[23,1169,362],{"id":361},[364,1171,1173],{"id":1172},"what-is-the-best-free-firecrawl-alternative","What is the best free Firecrawl alternative?",[11,1175,1176],{},"Crawl4AI self-hosted, or Jina Reader for zero-setup single pages.",[364,1178,1180],{"id":1179},"is-firecrawl-worth-it-vs-open-source","Is Firecrawl worth it vs open source?",[11,1182,1183],{},"Below serious volume, usually yes - $16/mo for 5,000 pages is cheaper than maintaining a crawler stack. OSS wins on volume and data control.",[364,1185,1187],{"id":1186},"what-cant-firecrawl-extract","What can't Firecrawl extract?",[11,1189,1190],{},"The design layer - tokens, brand, typography, assets, component code. Markdown strips it by design.",[364,1192,1194],{"id":1193},"which-alternative-is-best-for-ai-agents","Which alternative is best for AI agents?",[11,1196,1197],{},"Research agents: Tavily. Page-reading agents: Jina or Firecrawl. UI-building agents: MiroMiro.",[398,1199,400],{},{"title":152,"searchDepth":181,"depth":181,"links":1201},[1202,1203,1204,1205,1206,1207,1208,1209],{"id":897,"depth":181,"text":898},{"id":927,"depth":181,"text":928},{"id":960,"depth":181,"text":961},{"id":993,"depth":181,"text":994},{"id":1022,"depth":181,"text":1023},{"id":1086,"depth":181,"text":1087},{"id":344,"depth":181,"text":345},{"id":361,"depth":181,"text":362,"children":1210},[1211,1212,1213,1214],{"id":1172,"depth":412,"text":1173},{"id":1179,"depth":412,"text":1180},{"id":1186,"depth":412,"text":1187},{"id":1193,"depth":412,"text":1194},"Five real Firecrawl alternatives in 2026, matched to the job: cheap markdown, open source, site automation, agent search, or design context for AI.",[1217,1219,1222,1225],{"question":1173,"answer":1218},"Crawl4AI if you can self-host - it is open source, actively developed, and produces LLM-ready markdown with no per-page cost beyond your infrastructure. Jina Reader's free tier is the no-setup option: prefix any URL with r.jina.ai and get markdown back.",{"question":1220,"answer":1221},"Is Firecrawl worth it compared to open source scrapers?","If your time has a price, usually yes. Self-hosting a crawler means owning proxies, JS rendering, retries, and anti-bot upkeep. Firecrawl's Hobby tier at $16/mo for 5,000 pages is cheaper than the first afternoon of maintaining that stack. Open source wins on volume and data control, not convenience.",{"question":1223,"answer":1224},"What can't Firecrawl extract from a website?","The design. Firecrawl converts pages to markdown and structured text for LLMs - by design, that strips the visual layer. Brand colors, typography, spacing, design tokens, SVGs, and component code need a design-extraction tool; no volume of markdown credits gets you there.",{"question":1226,"answer":1227},"Which Firecrawl alternative is best for AI agents?","Depends on the loop. For research agents that answer questions, Tavily's search-shaped API fits better than crawling. For agents that read specific pages, Jina Reader or Firecrawl itself. For agents that build UI and need a site's design system as ground truth, a design-extraction API like MiroMiro.",{},"/api-blog/best-firecrawl-alternatives-2026",[876,1231,437],"best-brand-data-apis-2026",{"title":886,"description":1215},"api-blog/best-firecrawl-alternatives-2026",[1235,1236,1237,1238,446],"firecrawl","alternatives","web scraping","ai agents","IwAwF9vtdiqk_AlwZsEggBXVUHTcQNd0fz_Wwe13QZU",{"id":1241,"title":1242,"author":6,"body":1243,"category":1700,"date":417,"dateModified":418,"description":1701,"extension":420,"faqs":1702,"image":418,"meta":1714,"navigation":433,"path":1715,"readingTime":435,"relatedSlugs":1716,"seo":1718,"stem":1719,"tags":1720,"__hash__":1725},"apiBlog/api-blog/how-to-extract-brand-assets-in-bulk.md","How to Extract Brand Assets from 1,000 URLs in Bulk (Python)",{"type":8,"value":1244,"toc":1687},[1245,1248,1252,1263,1267,1283,1535,1545,1549,1568,1574,1584,1588,1637,1640,1644,1651,1653,1657,1660,1664,1667,1671,1674,1678,1684],[11,1246,1247],{},"The person with 1,000 URLs is never the person who wants to babysit 1,000 browser tabs. Bulk brand extraction in Python is an asyncio problem wearing a design-data hat: a semaphore for rate limits, retries for the 10% of URLs that misbehave, and checkpointing so a crash at URL 900 does not cost you the first 899. Here is the working script, then the failure math.",[23,1249,1251],{"id":1250},"why-not-1000-playwright-launches","Why not 1,000 Playwright launches",[11,1253,1254,1255,1258,1259,1262],{},"The ",[30,1256,1257],{"href":557},"single-site Playwright approach"," does not scale linearly. A thousand headless Chromium sessions means gigabytes of RAM, cookie banners in six languages, and a long tail of sites that hang ",[15,1260,1261],{},"networkidle"," forever. It is buildable - a browser pool, per-site timeouts, crash recovery - but at that point you are maintaining scraping infrastructure as a hobby. For bulk, put the browser problem on someone else's side of the API.",[23,1264,1266],{"id":1265},"the-script","The script",[147,1268,1270],{"className":149,"code":1269,"language":151,"meta":152,"style":152},"pip install httpx\n",[15,1271,1272],{"__ignoreMap":152},[156,1273,1274,1277,1280],{"class":158,"line":159},[156,1275,1276],{"class":162},"pip",[156,1278,1279],{"class":170}," install",[156,1281,1282],{"class":170}," httpx\n",[147,1284,1288],{"className":1285,"code":1286,"language":1287,"meta":152,"style":152},"language-python shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","import asyncio, csv, json\nimport httpx\n\nAPI = \"https://miromiro.app/api/v1/brand\"\nKEY = \"mm_live_your_key\"\nCONCURRENCY = 5          # stay under your plan's req/min cap\nRETRIES = 3\n\nasync def extract(client, sem, url):\n    async with sem:\n        for attempt in range(RETRIES):\n            try:\n                r = await client.get(\n                    API,\n                    params={\"url\": url, \"fields\": \"logo,palette\"},\n                    headers={\"Authorization\": f\"Bearer {KEY}\"},\n                    timeout=60,\n                )\n                if r.status_code == 429:            # rate limited: back off\n                    await asyncio.sleep(2 ** attempt * 5)\n                    continue\n                if r.status_code == 200:\n                    return {\"url\": url, \"ok\": True, \"data\": r.json()}\n                return {\"url\": url, \"ok\": False, \"error\": f\"HTTP {r.status_code}\"}\n            except httpx.HTTPError as e:\n                if attempt == RETRIES - 1:\n                    return {\"url\": url, \"ok\": False, \"error\": str(e)}\n                await asyncio.sleep(2 ** attempt)\n\nasync def main(urls):\n    sem = asyncio.Semaphore(CONCURRENCY)\n    async with httpx.AsyncClient() as client:\n        results = await asyncio.gather(*(extract(client, sem, u) for u in urls))\n    with open(\"brands.jsonl\", \"w\") as f:\n        for row in results:\n            f.write(json.dumps(row) + \"\\n\")\n    failed = [r for r in results if not r[\"ok\"]]\n    print(f\"{len(results) - len(failed)} ok, {len(failed)} failed\")\n\nif __name__ == \"__main__\":\n    urls = [row[0] for row in csv.reader(open(\"domains.csv\"))]\n    asyncio.run(main(urls))\n","python",[15,1289,1290,1295,1300,1305,1311,1317,1323,1328,1333,1339,1345,1351,1357,1363,1369,1375,1381,1387,1393,1399,1405,1411,1417,1423,1429,1435,1441,1447,1453,1458,1464,1470,1476,1482,1488,1494,1500,1506,1512,1517,1523,1529],{"__ignoreMap":152},[156,1291,1292],{"class":158,"line":159},[156,1293,1294],{},"import asyncio, csv, json\n",[156,1296,1297],{"class":158,"line":181},[156,1298,1299],{},"import httpx\n",[156,1301,1302],{"class":158,"line":412},[156,1303,1304],{"emptyLinePlaceholder":433},"\n",[156,1306,1308],{"class":158,"line":1307},4,[156,1309,1310],{},"API = \"https://miromiro.app/api/v1/brand\"\n",[156,1312,1314],{"class":158,"line":1313},5,[156,1315,1316],{},"KEY = \"mm_live_your_key\"\n",[156,1318,1320],{"class":158,"line":1319},6,[156,1321,1322],{},"CONCURRENCY = 5          # stay under your plan's req/min cap\n",[156,1324,1325],{"class":158,"line":435},[156,1326,1327],{},"RETRIES = 3\n",[156,1329,1331],{"class":158,"line":1330},8,[156,1332,1304],{"emptyLinePlaceholder":433},[156,1334,1336],{"class":158,"line":1335},9,[156,1337,1338],{},"async def extract(client, sem, url):\n",[156,1340,1342],{"class":158,"line":1341},10,[156,1343,1344],{},"    async with sem:\n",[156,1346,1348],{"class":158,"line":1347},11,[156,1349,1350],{},"        for attempt in range(RETRIES):\n",[156,1352,1354],{"class":158,"line":1353},12,[156,1355,1356],{},"            try:\n",[156,1358,1360],{"class":158,"line":1359},13,[156,1361,1362],{},"                r = await client.get(\n",[156,1364,1366],{"class":158,"line":1365},14,[156,1367,1368],{},"                    API,\n",[156,1370,1372],{"class":158,"line":1371},15,[156,1373,1374],{},"                    params={\"url\": url, \"fields\": \"logo,palette\"},\n",[156,1376,1378],{"class":158,"line":1377},16,[156,1379,1380],{},"                    headers={\"Authorization\": f\"Bearer {KEY}\"},\n",[156,1382,1384],{"class":158,"line":1383},17,[156,1385,1386],{},"                    timeout=60,\n",[156,1388,1390],{"class":158,"line":1389},18,[156,1391,1392],{},"                )\n",[156,1394,1396],{"class":158,"line":1395},19,[156,1397,1398],{},"                if r.status_code == 429:            # rate limited: back off\n",[156,1400,1402],{"class":158,"line":1401},20,[156,1403,1404],{},"                    await asyncio.sleep(2 ** attempt * 5)\n",[156,1406,1408],{"class":158,"line":1407},21,[156,1409,1410],{},"                    continue\n",[156,1412,1414],{"class":158,"line":1413},22,[156,1415,1416],{},"                if r.status_code == 200:\n",[156,1418,1420],{"class":158,"line":1419},23,[156,1421,1422],{},"                    return {\"url\": url, \"ok\": True, \"data\": r.json()}\n",[156,1424,1426],{"class":158,"line":1425},24,[156,1427,1428],{},"                return {\"url\": url, \"ok\": False, \"error\": f\"HTTP {r.status_code}\"}\n",[156,1430,1432],{"class":158,"line":1431},25,[156,1433,1434],{},"            except httpx.HTTPError as e:\n",[156,1436,1438],{"class":158,"line":1437},26,[156,1439,1440],{},"                if attempt == RETRIES - 1:\n",[156,1442,1444],{"class":158,"line":1443},27,[156,1445,1446],{},"                    return {\"url\": url, \"ok\": False, \"error\": str(e)}\n",[156,1448,1450],{"class":158,"line":1449},28,[156,1451,1452],{},"                await asyncio.sleep(2 ** attempt)\n",[156,1454,1456],{"class":158,"line":1455},29,[156,1457,1304],{"emptyLinePlaceholder":433},[156,1459,1461],{"class":158,"line":1460},30,[156,1462,1463],{},"async def main(urls):\n",[156,1465,1467],{"class":158,"line":1466},31,[156,1468,1469],{},"    sem = asyncio.Semaphore(CONCURRENCY)\n",[156,1471,1473],{"class":158,"line":1472},32,[156,1474,1475],{},"    async with httpx.AsyncClient() as client:\n",[156,1477,1479],{"class":158,"line":1478},33,[156,1480,1481],{},"        results = await asyncio.gather(*(extract(client, sem, u) for u in urls))\n",[156,1483,1485],{"class":158,"line":1484},34,[156,1486,1487],{},"    with open(\"brands.jsonl\", \"w\") as f:\n",[156,1489,1491],{"class":158,"line":1490},35,[156,1492,1493],{},"        for row in results:\n",[156,1495,1497],{"class":158,"line":1496},36,[156,1498,1499],{},"            f.write(json.dumps(row) + \"\\n\")\n",[156,1501,1503],{"class":158,"line":1502},37,[156,1504,1505],{},"    failed = [r for r in results if not r[\"ok\"]]\n",[156,1507,1509],{"class":158,"line":1508},38,[156,1510,1511],{},"    print(f\"{len(results) - len(failed)} ok, {len(failed)} failed\")\n",[156,1513,1515],{"class":158,"line":1514},39,[156,1516,1304],{"emptyLinePlaceholder":433},[156,1518,1520],{"class":158,"line":1519},40,[156,1521,1522],{},"if __name__ == \"__main__\":\n",[156,1524,1526],{"class":158,"line":1525},41,[156,1527,1528],{},"    urls = [row[0] for row in csv.reader(open(\"domains.csv\"))]\n",[156,1530,1532],{"class":158,"line":1531},42,[156,1533,1534],{},"    asyncio.run(main(urls))\n",[11,1536,1537,1538,1540,1541,1544],{},"Swap the endpoint for ",[15,1539,139],{}," if you want full token sets instead of brand data, and adjust ",[15,1542,1543],{},"fields="," to trim the response to what you store - it does not change the credit cost, but it makes the output files dramatically smaller.",[23,1546,1548],{"id":1547},"the-three-things-that-make-bulk-runs-survive","The three things that make bulk runs survive",[11,1550,1551,1554,1555,1559,1560,1563,1564,1567],{},[41,1552,1553],{},"1. The semaphore is the contract."," Your plan has a requests-per-minute cap (the ladder is in the ",[30,1556,1558],{"href":1557},"/api/docs/rate-limits","rate limits docs","). Five concurrent with backoff on 429 stays comfortably inside the entry tiers; raise it as your plan allows. Blasting unbounded ",[15,1561,1562],{},"gather"," at any API gets you rate-limited into taking ",[207,1565,1566],{},"longer"," than the polite version.",[11,1569,1570,1573],{},[41,1571,1572],{},"2. Expect 5-15% failures and capture them."," On a real 1,000-domain list some are dead, parked, or behind bot walls. The script writes per-URL errors instead of dying, so the run always finishes and the failures are a re-runnable list, not a mystery.",[11,1575,1576,1579,1580,1583],{},[41,1577,1578],{},"3. Re-runs are cheap because of caching."," Responses are cached for 24 hours and a cache hit costs 0 credits (",[15,1581,1582],{},"\"cached\": true"," in the usage object). Re-running today's failed batch only pays for URLs that actually re-extract - so retry freely.",[23,1585,1587],{"id":1586},"budgeting-a-1000-url-run","Budgeting a 1,000-URL run",[226,1589,1590,1602],{},[229,1591,1592],{},[232,1593,1594,1596,1599],{},[235,1595],{},[235,1597,1598],{},"Brand data (15 cr/URL)",[235,1600,1601],{},"Design tokens (10 cr/URL)",[257,1603,1604,1615,1626],{},[232,1605,1606,1609,1612],{},[262,1607,1608],{},"100 URLs",[262,1610,1611],{},"1,500 credits",[262,1613,1614],{},"1,000 credits",[232,1616,1617,1620,1623],{},[262,1618,1619],{},"1,000 URLs",[262,1621,1622],{},"15,000 credits",[262,1624,1625],{},"10,000 credits",[232,1627,1628,1631,1634],{},[262,1629,1630],{},"Fits in",[262,1632,1633],{},"Growth (€79/mo, 30k)",[262,1635,1636],{},"Growth, or 2× Developer months",[11,1638,1639],{},"The free 100 credits cover a 6-10 URL pilot - do that first, on your ugliest URLs, before committing a plan to the full list. Wall-clock: the rate limiter sets the pace, not the extraction; expect tens of minutes for 1,000 URLs, not hours.",[23,1641,1643],{"id":1642},"summary","Summary",[11,1645,1646,1647,1650],{},"Bulk extraction is a queue-discipline problem: bounded concurrency, backoff on 429, per-URL error capture, and a re-run list. The 24-hour cache turns retries from a cost into a free operation. Pilot on 10 URLs with a ",[30,1648,1649],{"href":351},"free key",", then size the plan to the list.",[23,1652,362],{"id":361},[364,1654,1656],{"id":1655},"how-do-i-scrape-brand-colors-from-a-list-of-websites-in-python","How do I scrape brand colors from a list of websites in Python?",[11,1658,1659],{},"For a few sites, Playwright and computed styles. For hundreds or thousands, the async httpx script above against an extraction API.",[364,1661,1663],{"id":1662},"how-long-does-1000-urls-take","How long does 1,000 URLs take?",[11,1665,1666],{},"Tens of minutes - your requests-per-minute cap is the bottleneck, not the extraction itself.",[364,1668,1670],{"id":1669},"what-failure-rate-should-i-expect","What failure rate should I expect?",[11,1672,1673],{},"5-15% on real-world lists: dead domains, parked pages, bot walls. Capture and re-run rather than aiming for zero.",[364,1675,1677],{"id":1676},"do-repeated-extractions-cost-credits-again","Do repeated extractions cost credits again?",[11,1679,1680,1681,559],{},"Not within 24 hours - cache hits cost 0 credits and report ",[15,1682,1683],{},"cached: true",[398,1685,1686],{},"html pre.shiki code .sBMFI, html code.shiki .sBMFI{--shiki-light:#E2931D;--shiki-default:#FFCB6B;--shiki-dark:#FFCB6B}html pre.shiki code .sfazB, html code.shiki .sfazB{--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D}html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}",{"title":152,"searchDepth":181,"depth":181,"links":1688},[1689,1690,1691,1692,1693,1694],{"id":1250,"depth":181,"text":1251},{"id":1265,"depth":181,"text":1266},{"id":1547,"depth":181,"text":1548},{"id":1586,"depth":181,"text":1587},{"id":1642,"depth":181,"text":1643},{"id":361,"depth":181,"text":362,"children":1695},[1696,1697,1698,1699],{"id":1655,"depth":412,"text":1656},{"id":1662,"depth":412,"text":1663},{"id":1669,"depth":412,"text":1670},{"id":1676,"depth":412,"text":1677},"Tutorials","Bulk-extract brand data and design tokens from thousands of domains in Python - async code with rate limiting and retries that survives a real URL list.",[1703,1705,1708,1711],{"question":1656,"answer":1704},"For a handful of sites, a Playwright script reading computed styles works. Past a few hundred URLs, run the extraction through an API with asyncio and httpx: a semaphore to respect rate limits, retries with backoff for the URLs that fail, and checkpointing so a crash does not restart the whole run.",{"question":1706,"answer":1707},"How long does it take to extract data from 1,000 URLs?","With an API at a typical plan rate limit, roughly 20-60 minutes for 1,000 URLs depending on your requests-per-minute cap - the limiter, not the extraction, sets the pace. A self-hosted browser farm is usually slower per URL once retries and crashes are counted.",{"question":1709,"answer":1710},"What percentage of URLs fail in a bulk extraction run?","Plan for 5-15% on a real-world list: dead domains, redirects to parked pages, bot walls, timeouts. This is why retries and per-URL error capture matter more than raw speed - a bulk script that stops on the first 404 never finishes.",{"question":1712,"answer":1713},"Do repeated extractions of the same URL cost credits again?","On MiroMiro, responses are cached for 24 hours and a cache hit costs 0 credits - the usage object reports cached: true. Re-running a failed batch the same day only pays for the URLs that actually re-extract.",{},"/api-blog/how-to-extract-brand-assets-in-bulk",[1231,874,1717],"add-design-extraction-to-claude-code-mcp",{"title":1242,"description":1701},"api-blog/how-to-extract-brand-assets-in-bulk",[1721,1287,1722,1723,1724],"bulk extraction","async","brand data","automation","s_wWV-k0k55QCGaVhQYXcQuGKhALy3iIWTRlMgsjUVM",1784751617348]