Scraping, document processing, programmatic publishing — pipelines built by hand, scheduled and monitored, handed over with the keys. The receipts on this page are from systems I actually run.
What it runs
Scrapers and curators that built a 257,918-article corpus across 50+ sources — cleaned, deduplicated, and ready for retrieval or training.
Extraction, comparison, entity work — the machinery behind nerscope, a pip-installable NLP library with a full test suite and CI.
Reports, catalogues, and print-ready documents generated from data — like the investigative zine built as a ReportLab pipeline.
Receipts
How it goes
A free call. You describe the weekly grind; I map inputs, outputs, and edge cases — then a fixed written quote within 48 hours.
Hand-written Python, tested on your real data, with logs and alerts so silence never means failure.
Scheduled, documented, yours. Your team owns it; I stay reachable.