A healthcare research firm in the Netherlands needed pharma competitive intelligence spanning every medicine category on leading pharmaceutical portals, sustained across a 3 year analysis window.
Client Overview
A healthcare research firm based in the Netherlands needed pharma competitive intelligence covering the full catalog of leading pharmaceutical portals, not a narrow slice of medicines but every category those portals carried. The firm planned to run its analysis over a three year window, which meant the data pipeline behind it needed to hold up as a sustained, ongoing source rather than a single extraction project.
That timeframe changed the nature of the problem. A three year study needs consistent, comparable data collected the same way month after month, not a one-off pull that happens to be thorough. Getting this kind of data at scale, across every category, sustained over years rather than weeks, is what brought the firm to PromptCloud.
Client Requirements
The firm’s brief to PromptCloud was specific about both breadth and depth:
- Medicine data across all categories available on the target pharmaceutical portals
- Regular, recurring extraction sustained over a three year analysis window
- Disease or condition name for every medicine record
- Manufacturer name and salt name captured alongside each entry
- Medicine name and price, kept current as portals updated their catalogs
Challenges
Covering every category on a pharmaceutical portal, rather than a single therapeutic area, meant the crawl could not be tuned narrowly for one kind of medicine listing. Disease and condition names, manufacturer details, salt names, medicine names, and prices all needed to be captured consistently across a catalog broad enough to support genuine competitive intelligence rather than a partial view of the market.
The three year timeframe was the harder problem. Big data volume was inevitable given the category breadth and the sustained collection window, and the firm needed real infrastructure behind this, not a script that would need constant babysitting to keep running reliably for years rather than weeks.
Solutions
PromptCloud built a custom crawler sized for both the category breadth and the multi-year timeline, backed by monitoring and a queryable API rather than a one-time data dump.
A Custom Crawler Built for the Full Catalog
Rather than adapting a generic scraper, PromptCloud built a custom crawler specifically for the structure of the target pharmaceutical portals, covering every category rather than a narrow subset. This kind of pharma competitive intelligence depends on catalog completeness as much as data accuracy, since missing categories would have left real gaps in the firm’s three year analysis. The same underlying approach connects directly to PromptCloud’s broader work in pricing intelligence, since price was one of the five core fields tracked for every medicine record.
Five Fields Captured for Every Medicine
Every record extracted carried disease or condition name, manufacturer name, salt name, medicine name, and price, the specific fields the firm needed to actually compare medicines across manufacturers and categories. Capturing all five consistently, rather than treating price or manufacturer as optional extras, is what made the resulting data usable for real competitive analysis instead of a simple product listing. That consistency mattered even more given the data needed to remain comparable across a three year window.
Monitoring Built for a Multi-Year Timeline
Because this data collection was meant to run for three years, not three months, PromptCloud set up monitoring on the target portals to catch structural changes and update the crawler promptly whenever a site changed. A crawler that silently breaks partway through a multi-year study creates a gap in the data that undermines the analysis built on top of it, so ongoing monitoring was treated as core infrastructure rather than an afterthought.
Delivered Through an API, Not a One-Time Export
The firm was given API access and documentation to query and fetch crawled data directly, rather than receiving periodic flat file exports. That approach suited a three year engagement far better than a recurring manual handoff would have, letting the firm’s own systems pull fresh pharma competitive intelligence on its own schedule as the study progressed.
Pharma Data Collection, Before and After PromptCloud
| Area | Before | After |
| Category coverage | Not addressed by a narrow, single-category scrape | Full catalog coverage across all portal categories |
| Data consistency | Risk of gaps over a multi-year window | Five core fields captured consistently for every record |
| Site changes | Would silently break a long-running crawl | Monitored and updated promptly as portals change |
| Data access | Would need repeated manual exports | Queryable via API and documentation |
Benefits to the Client
The firm got pharma competitive intelligence covering every category on the target portals, not a partial view limited to one therapeutic area, giving its three year study a genuinely comprehensive base to work from. Five consistent fields, disease or condition, manufacturer, salt name, medicine name, and price, meant every record supported real cross-manufacturer comparison rather than a bare product listing.
Monitoring built specifically for a multi-year timeline meant the firm’s data collection kept running reliably as target portals changed their own site structures over time, without gaps the firm would have had to notice and chase down itself. API access meant the firm could pull fresh data on its own schedule throughout the three year window rather than waiting on repeated manual deliveries.
Pharma Competitive Intelligence That Holds Up Over Three Years, Not Three Months
A study built to run for three years cannot depend on a crawler tuned for a single extraction. Pharma competitive intelligence at this scale needs full category coverage, consistent fields across every record, and monitoring built to catch site changes long before they turn into gaps the analysis cannot recover from.
That is what turned a broad, multi-year data requirement into something the firm could actually build a longitudinal study on, comprehensive from the first month to the last.



