Data, extracted — data, under control
Everything built here is data work. One half goes out and gets it: scrapers that pull structured records off any site, at scale. The other half deals with the data you already hold — the registers, inventories, assessments and dashboards that keep it organised, scored and auditable.
390 production scrapers on Apify, billed per result. 58 Excel and Google Sheets tools on Etsy for information security, data protection, risk and operations. Same subject, both ends of it.
Most-run scrapers
Structured data from real-world sites, billed per result on Apify. Ranked by total runs.
Website Contact Scraper – Email, Phone & Social
Extract emails, phone numbers and social links from any website in bulk. A no-API contact scraper that exports leads to CSV or JSON for outreach and CRMs.
Reddit Historical Archive Scraper - Old Posts by Date
Scrape old Reddit posts and comments by date as a Pushshift alternative. Full-text comment search and user history exported to CSV or JSON. No API key or login.
GMGN Scraper — Trending Memecoins & Smart Money
Scrape GMGN.ai without login: export trending memecoins, smart-money wallets, new pairs and rugchecks to CSV or JSON. Unofficial GMGN API alternative, no key.
Bulk URL Status Checker – Broken Link & Redirect Audit
Check HTTP status codes for thousands of URLs at once: find 404 broken links, trace 301/302 redirect chains and response times. Export results to CSV or JSON.
Threads Scraper - Scrape threads.net Posts, No Login
Scrape Threads posts, profiles and replies by username, URL or keyword. No-login Threads API alternative to export public post and engagement data to CSV or JSON.
Welcome to the Jungle Jobs Scraper – WTTJ Data
Scrape Welcome to the Jungle (WTTJ) jobs without login. Unofficial API alternative for salary and company data export to CSV or JSON — English and French listings.
Getting data out of a hostile website and keeping a company's own data in order are the same discipline wearing different clothes: decide what a record is, define the fields, validate what goes in them, and make the result something you can query and defend. One job ends in a dataset, the other in a register an auditor can read. The thinking does not change.
Excel & Google Sheets data tools
Information-security, privacy, risk, audit and operations registers — with the scoring, validation and dashboards already built. $14.90–$69, on Etsy.
ISO 27001 & information security
The working files an ISMS actually runs on — risk register, Statement of Applicability, internal audit, asset inventory, access reviews, incident log. Scoring, status and dashboards are already wired; you fill in your organisation, not the formulas.
Data protection & privacy (GDPR)
Everything a privacy programme has to be able to show on request: a record of processing, DPIA screening and scoring, data subject requests with the clock running, breach log, retention schedule. Built to be filled in, not read.
Frameworks & regulatory readiness
Self-assessment and readiness trackers for the frameworks that show up in vendor questionnaires and contracts — each one a control-by-control checklist with maturity scoring, gaps, owners and a readiness dashboard.
Quality, environment & safety
The management-system files that keep a certification alive between audits: clause 4–10 gap analysis, objectives, internal audit programme, nonconformities and corrective actions, supplier evaluation, competency and SOP control.
3 listed on Etsy so far · 3–5 new toolkits published a day.
Browse scrapers by category
19 categories, every actor tagged for fast discovery.
B2B contact, registry & prospect data for sales pipelines.
Product catalogs, prices, merchant intelligence.
APIs, datasets and dev infrastructure scrapers.
Competitor intel, market research, registries.
Property listings, prices, agents — EU & global.
Reddit, LinkedIn, podcast & content platforms.
Job board scrapers across geos and verticals.
Headless workflows, data pipelines, API replacements.
Ad libraries, campaign data, brand monitoring.
News aggregation, RSS, content monitoring.
Video & podcast platform data.
Sitemap, schema, broken link, technical SEO data.
Niche datasets and utilities.
Hotel prices, OTA, destination data.
AI training data, models, datasets, RAG inputs.
Game stores, catalogues, reviews and player data.
Routes, scores, live event data.
Tools built to be called by autonomous agents.
Model Context Protocol endpoints for LLM clients.
Latest guides
Deep dives on getting real-world sites to give up their data.
How to Build a Deep-Research Retrieval Layer for AI Agents
Give an agent a topic and get ranked web sources, full-page Markdown, and recent news with citations — keyless multi-source retrieval for RAG and grounding.
How to Extract Structured Data from Any URL in 2026
Turn any web page into clean JSON without an LLM: parse schema.org JSON-LD, OpenGraph, tables, prices and contacts deterministically — no API key, no browser.
How to Add Live Web Search to AI Agents (No API Key)
Give your LLM fresh, citable search results without a Tavily or SerpAPI key. How live SERP extraction works: ranked results plus Markdown page content for RAG.
How to Scrape Allabolag.se Sweden Company Leads in 2026
Extract Swedish company data from allabolag.se — org number, revenue, CEO, phone and email — without an API key. A guide to bulk firmographics and B2B leads.
How to Scrape arXiv Papers, Abstracts & Metadata in 2026
Build a research-paper dataset from arXiv's public API: titles, full abstracts, author lists, categories, PDF links and DOIs. No API key, no browser, at scale.
How to Scrape the Australia Business Register (ABN/ABR)
Bulk-export Australian business names, ABNs, status and registration dates from the official data.gov.au CKAN register — no API key, no browser, no captcha.
Got a data problem that isn't covered yet?
Data you need pulled out of a site, or data you keep wrangling by hand in a spreadsheet that was never designed for it. Either end, say what it is and it gets built.