Efath Web Scraper
Cold outreach lives or dies on the first line.
Cold outreach lives or dies on the first line. Every sales email that opens with "I hope this finds you well" gets deleted before the second sentence, and everyone doing outbound sales knows it, and almost nobody does anything about it, because writing a genuinely personalized opener for every single prospect on a list is slow, tedious work that doesn't scale past the first dozen names. This is a small internal tool built to solve exactly that: automate the research, not the pitch. Feed it a company's homepage, get back one real sentence, in the target language, that references something actually true about that company, not a mail-merge token, an actual sentence a human would have to have visited the site to write.
Built specifically for Dutch B2B outreach (the target list is a handful of real Dutch companies), the pipeline visits each homepage, scrapes what's actually there, and hands the extracted content to OpenAI with a tightly constrained system prompt: exactly one Dutch sentence, capped at thirty-five words, required to reference something real from the page, explicitly forbidden from inventing facts that weren't actually on the site. That last constraint mattered the most to get right. A cold email opener that name-drops something false about the prospect's company is worse than no personalization at all; it reads as either lazy or dishonest the moment the recipient notices, and either read kills the email instantly.
The pipeline, mechanically
Python, requests plus BeautifulSoup for the fast path, with an automatic fallback to Playwright running headless Chromium for anything JavaScript-heavy or actively blocking simple scrapers. That fallback matters more than it might sound like: plenty of modern corporate sites render their real content client-side, and a scraper that only handles static HTML would come back empty-handed on exactly the sites most worth researching. Results land in a CSV, every run logs to a rotating log file, and the whole thing is written to be fully portable. Every path is relative to the project folder itself, not to whatever directory you happened to run it from, specifically so the tool can be zipped up and handed to a different machine without breaking.
There's a comment block in main.py that reads like a real production checklist, not a to-do list I never finished: SSL verification actually enabled, retry-with-backoff actually implemented, no magic numbers scattered through the logic, full type hints throughout. Somewhere around twenty-five checked items, all genuinely addressed in the code, not just listed and forgotten. For a small internal tool that nobody outside this project will ever see, that's a level of discipline I didn't strictly need to apply. Nobody was going to audit this code for production readiness, and I applied it anyway, because sloppy code has a way of costing you later regardless of how small the tool started out, and because building things correctly the first time is a habit I don't want to selectively turn off just because a project feels low-stakes.
The honest result of an actual run
output.csv is a real run's output, not a demo file I cleaned up for presentation. Five target companies. Two succeeded: Picnic and Exact both came back with genuine, usable Dutch sentences. Three came back marked ERROR, Mollie, Transavia, and Vandebron, almost certainly because their sites either blocked the scraper outright or presented something the extraction logic couldn't parse cleanly. A forty percent success rate on a five-company sample isn't a number I'd want printed on a sales deck, and I'm not going to pretend otherwise. It's also an honest snapshot of what real-world scraping actually looks like the first time you point it at real, defended corporate websites instead of a controlled test environment. Some sites cooperate, some don't, and building a tool that fails loudly and specifically on the ones that don't, instead of silently returning garbage, is worth more than a tool that claims a hundred percent success rate by quietly skipping the hard cases.
The CSV writes a row for every company regardless of outcome, success or failure, which was a deliberate choice. You always know exactly which prospects still need a human to write their opener by hand, instead of a report that only shows you the wins and leaves the failures invisible until someone notices a name missing.
Why sequential, not parallel
The scraping runs strictly one site at a time, no threading, no async batching. That's slower than it needs to be on raw throughput, and I built it that way anyway, because the whole point of scraping five specific company homepages for a personalized outreach campaign is not the same problem as scraping five thousand pages for a dataset. Rate-limit safety and not looking like an aggressive bot matters more here than speed. A scraper hammering a handful of real companies' sites in parallel is more likely to get itself blocked than one making polite, sequential, spaced-out requests. For this specific use case, slow and reliable beats fast and flagged.
What it actually is, scoped honestly
Five companies is a toy-scale list, not a production lead-gen pipeline. This was built to prove the pattern: scrape, extract, generate one honest sentence, log everything, fail visibly instead of silently, on a small real sample before deciding whether it was worth scaling to an actual prospect list in the hundreds or thousands. It's a tool built to answer "does this approach actually work" before committing more time to "does this approach scale," which is the right order to build things in even when it means the first version looks small next to what it's eventually meant to become.
What the three failures actually taught me
I looked at the three errored companies more closely than I initially planned to, because a forty percent failure rate on a five-company list is small enough to actually investigate by hand instead of just shrugging at an aggregate number. Mollie, Transavia, and Vandebron are all larger, more established companies than Picnic or Exact, and larger companies tend to run more aggressive bot-detection on their public-facing sites, which tracks with what actually happened here. The tool wasn't broken. It was correctly, honestly reporting that it got blocked, rather than returning something that looked like success while quietly scraping garbage. I'd rather have a scraper that fails loudly on the sites that actively resist it than one that silently hands back malformed data and lets a sales rep send a broken, nonsensical opener to a real prospect without knowing it.
The system prompt as the actual product
If I had to point at the single most important piece of this whole tool, it wouldn't be the scraping logic. Scraping a webpage is a well-understood problem with well-understood solutions. It would be the system prompt constraining what OpenAI is allowed to generate from what gets scraped: one sentence, thirty-five words, Dutch, grounded in something real from the page, no invented facts. Getting that constraint tight enough to consistently produce something usable, without either hallucinating specifics that aren't there or producing something so generic it might as well be a template, took more iteration than the scraping code did. The scraper gets you the raw material. The prompt is what turns raw material into something a real salesperson could actually paste into an email without embarrassing themselves.
Full stack developer. Founder of Yashveer Labs. One real sentence, honestly earned, beats a hundred generic ones.
Start a conversation about this.
Whether it's efath web scraper itself or the next system worth building, the lab is reachable.