A US ecommerce platform needed web scraping for ecommerce product data, including every color and size variant, from more than 100 fashion retailers to launch its own fashion marketplace app.
Client Overview
A fast growing ecommerce platform in the United States set out to build a fashion marketplace app that could show shoppers products from across the retail landscape in one place, rather than sending them to a hundred different sites to compare options. The idea only worked if the app had real product data behind it, names, descriptions, prices, and every color and size variant a shopper might want to check.
Building that catalog meant pulling data from more than 100 niche fashion retailers, including well known names like GAP, Macy’s, and Nordstrom, each running its own site with its own layout. Web scraping for ecommerce at this scale needed more than a handful of custom scripts, so the platform brought in PromptCloud to build the pipeline.
Client Requirements
The client came to PromptCloud with a clear list of what the marketplace app needed:
- Product data extraction across more than 100 fashion retailer websites
- Full variant coverage, meaning every color and size option per product
- A daily extraction schedule to keep the catalog current
- Delivery in CSV format, uploaded directly to the client’s S3 servers
- A source list and target data points defined upfront by the client’s own team
Challenges
Pulling product data from more than 100 independent retailer sites is a different problem than scraping one or two. Each site used its own layout, its own way of listing color and size options, and its own pagination and category structure, none of which could be standardized before the crawl began.
Variant data added another layer of complexity. A single product page could hide half a dozen size and color combinations behind dropdowns or swatches, and missing even one meant an incomplete listing once it reached the marketplace app. Getting all of it, consistently, across a source list this size, meant the crawl setup had to be built site by site rather than with one generic template.
Solutions
PromptCloud treated this as a site crawling project rather than a single template, the kind of web scraping for ecommerce that only works when each source gets its own logic, building extraction, delivery, and monitoring around the reality of 100-plus independent sources.
Site-by-Site Crawler Setup for 100+ Sources
Rather than forcing every retailer site through one generic scraper, PromptCloud’s team built crawlers tuned to each source, since GAP, Macy’s, Nordstrom, and the rest of the client’s list of 100-plus fashion sites each organize product listings differently. This site crawling approach handled the variation in layout and pagination without asking the client to normalize anything on their end. The source list and required data points came directly from the client, so the crawl setup was defined by what the marketplace app actually needed, not a generic fashion data template.
Capturing Every Color and Size Variant
Product data extraction only mattered if it was complete, so the crawlers were built to capture every size and color variant tied to a product, not just the default listing a shopper sees first. That meant reading dropdowns, swatches, and variant specific pricing and discounts on each page, then attaching all of it back to the parent product record. A marketplace app showing a product without its full variant set is showing shoppers an incomplete option, and this was the piece most likely to get missed by a simpler crawl.
Daily Delivery Straight to S3
Once extracted, data landed in CSV format directly in the client’s S3 servers on a daily schedule, with no manual handoff step in between. The initial setup was live within a few days, and crawlers began delivering data as soon as it was ready. For teams weighing whether to build this kind of pipeline internally, PromptCloud’s DIY web scraping cost report lays out what that decision typically costs in engineering time alone.
Monitoring 100+ Sites for Change
A crawl setup this size does not stay static. Retailer sites redesign pages, change how variants are displayed, or restructure categories without warning, and any one of those changes can silently break a crawler tuned to the old layout. PromptCloud set up ongoing monitoring across the full source list so that changes on any of the 100-plus sites triggered a fix before the client noticed a gap in the data, rather than after.
Fashion Data Pipeline, Before and After PromptCloud
| Area | Before | After |
| Product coverage | No unified data across 100+ retailers | Standardized data from every source, daily |
| Variant data | Manual checking risked missed sizes and colors | Every color and size variant captured automatically |
| Delivery | No automated pipeline to S3 | Daily CSV delivery straight to S3, no manual step |
| Site changes | Would break crawlers silently | Monitored continuously, fixed before gaps appear |
Benefits to the Client
The marketplace app started receiving real product data within days of the initial setup, well before the client would have finished building an equivalent pipeline in house. PromptCloud’s team handled every technical piece of the extraction, from site specific crawler logic to variant capture to delivery, so the client’s own team could focus on the app rather than the data behind it.
Volume scaled quickly. The first crawl delivered about 200,000 records, and once the full 100-plus site list was running on its daily schedule, volume grew to more than 1 million records a day. That scale, handled without added headcount on the client’s side, let the marketplace app launch on a timeline that would have been difficult to hit building the pipeline from scratch.
Web Scraping for Ecommerce That Scales With the Catalog
A fashion marketplace built on data from 100-plus retailers cannot run on a handful of manually maintained scripts. Web scraping for ecommerce at this scale means treating each source as its own problem, capturing every variant a shopper might look for, and catching site changes before they turn into gaps in the catalog.
None of that changes the basic goal, get complete, current product data to the app reliably enough that the business built on top of it can launch and keep growing. That is what turned a hundred different retailer sites into one marketplace catalog the client could actually ship.



