How AI crawlers read the web
50 real websites across 9 industries, every request logged at the edge. Which fetchers visit, what they open first, how often they return, and what they ignore. This is the backbone of everything else the lab publishes.
Brainwork is the independent research lab run by Misha Manko. It operates an instrumented network of 50 real websites, records every visit from every AI crawler, measures which sources AI answers actually cite, and turns the data into published findings and working instruments.
No surveys of what people think AI does. No guesses. The lab runs on measurement.
Each program has its own instrument, its own data, and its own publication trail. None of them rely on vendor dashboards or third-party estimates.
50 real websites across 9 industries, every request logged at the edge. Which fetchers visit, what they open first, how often they return, and what they ignore. This is the backbone of everything else the lab publishes.
Common Crawl feeds most training corpora. The lab measures how much of a domain each crawl actually captured, when robots.txt started blocking it, how far the stored copy has drifted from the live page, and where the domain sits in the 118-million-node web graph.
Nightly runs of real buyer prompts through ChatGPT and Google AI Overviews, with every cited source stored raw. The question is not "are we mentioned" but "which pages earn citations, and what do they have in common."
Each one started as a measurement problem inside a research program. When the data needed a tool, the lab built it and kept it running.
Enter a domain, get seven panels: captures per crawl, robots.txt block history with dates, stored-copy drift, a live CCBot probe, sitemap coverage gaps, and the domain's crawl-priority rank in the full web graph.
LiveRuns a client's real buyer prompts through ChatGPT and Google AI Overviews every night, stores every raw response, and reports which sources were cited. Built after testing the commercial trackers and finding the numbers mostly noise.
In serviceA desktop widget that shows AI usage limits as rings in the corner of the screen. Built for the lab's own heavy use of coding agents, then productized for anyone who hits the same walls.
LaunchingEvery claim starts with a sensor. Real sites, real edge logs, real API responses stored raw. If it cannot be measured, the lab does not have an opinion on it yet.
Continuously, not as a one-off study. AI fetchers change behaviour month to month. A snapshot from last quarter is already a historical document.
Findings go out in the open with the method attached, including the ones that contradict the industry's favourite advice. The next person should not have to repeat the work.
Published research from the network, written up on Misha Manko's site. Method and data described in every piece.
Research collaborations, data questions, and consulting on the harder cases. One business day to reply. Productized audits and engagements run through mishamanko.com.