API
Public, read-only, no key. Counts of crawlers rather than of readers: this site stores no IP addresses at all, and keeps a raw User-Agent only for a request it classified as a bot. CORS is open, so you can call it straight from a browser.
Agents
GET https://willblew.com/api/agents/
Which crawlers and AI agents have been here, how much they took, and what each one appears to be doing.
| Parameter | Values | What it does |
|---|---|---|
window | 24h, 7d, 30d, 90d, all | How far back to look. Defaults to 7d. |
agent | an agent name | Narrows to one agent and adds the exact paths it asked for. |
series | 1 | Adds a per-day series of total, bot and AI hits. |
curl -s 'https://willblew.com/api/agents/?window=30d&series=1'
curl -s 'https://willblew.com/api/agents/?agent=GPTBot&window=all'
What comes back
{
"api": "agentwatch",
"window": { "label": "7d", "from": "...", "to": "..." },
"classifier": { "version": "2026-09-29", "known_agents": 95 },
"totals": { "hits": 0, "bot_hits": 0, "human_hits": 0,
"ai_hits": 0, "probes": 0, "agents": 0 },
"categories": [ { "category": "ai-training", "hits": 0, "agents": 0,
"description": "..." } ],
"agents": [
{
"agent": "GPTBot",
"operator": "OpenAI",
"category": "ai-training",
"hits": 0, "pages": 0, "posts": 0, "probes": 0,
"read_rules": 0, "loaded_assets": 0, "loaded_images": 0,
"days_active": 0,
"first_seen": "...", "last_seen": "...",
"suspected": {
"verdict": "Harvesting the archive for training",
"confidence": "high",
"evidence": [ "read robots.txt or llms.txt first",
"12 distinct posts", "across 4 days" ]
},
"user_agent": "..."
}
]
}
Categories
| Category | Meaning |
|---|---|
ai-training | Collects pages as material for training models. |
ai-search | Builds an index that an AI product answers from. |
ai-assistant | Fetches a page live because a person asked an assistant about it. |
search | Traditional search engine indexing. |
social | Unfurls a link into a preview card. |
seo | Harvests links and rankings for an SEO product. |
monitor | Uptime or availability checking. |
archive | Preserving pages for an archive. |
feed | Fetching a feed for a reader. |
tool | A generic HTTP client or scripting library. |
scanner | Looking for exposed files and known vulnerabilities. |
human | A browser driven by a person. |
unknown | Nothing in the User-Agent identifies it. |
How far to trust it
A User-Agent is a claim, not an identity. Anything here
can be forged by anyone who cares to, and plenty do. suspected is
the honest word: the verdict is a rule read off observed behaviour, and the
evidence that produced it is returned alongside so you can disagree with it.
The classifier is a data file with a version on it; when an operator renames
a crawler the file is what goes stale, and that version tells you how stale.