Skip to content

What is bot detection?

Bot detection is how a website decides whether a visitor is a person or an automated program. It reads signals on four layers: the network the request comes from, the browser that sends it, how the visitor behaves, and, when in doubt, a challenge such as a CAPTCHA. The result is usually a risk score, and the site decides per page whether to allow, slow down, challenge or block.

Sites use it to keep out credential stuffing, fake accounts, scalping, ad fraud and heavy scraping, while still letting in the bots they want, such as search engine crawlers, which can be verified. Most of it happens without the visitor noticing.

Modern detection is less about any single tell than about consistency: whether the TLS handshake, the user agent, the Client Hints, the fingerprint and the behaviour all describe one real browser. See how anti-bot systems detect automation and how detection works.

Updated

Good bots, bad bots, and the space in between

Not all automated traffic is unwanted. Search engine crawlers, uptime monitors, feed readers and link previews are bots that sites welcome. Credential stuffing, fake sign-ups, ticket and stock hoarding, ad fraud and aggressive scraping are bots that sites try to stop. A lot of automation sits in between: price monitoring, research crawlers, AI agents acting for a user, and testing tools.

Well-behaved crawlers identify themselves. A site can check a crawler's claim instead of trusting its user agent: Google, for example, says to run a reverse DNS lookup on the visiting IP address, confirm the name ends in googlebot.com, google.com or googleusercontent.com, and look the name up again to make sure it resolves to the same address, or to compare the address with the IP ranges Google publishes. The user agents of common crawlers and HTTP clients are listed with our user agent checker.

So bot detection is rarely a yes-or-no gate. Most systems produce a risk score and let the site decide per page: let the request through, slow it down, ask for a challenge, or block it.

How bot detection works: four layers

LayerWhat is checkedExamples
NetworkWhere the request comes from and how the connection was madeIP reputation and hosting ranges, request rate, the TLS and HTTP/2 fingerprint compared with the browser the request claims to be
BrowserWhat the client is, read by JavaScript in the pageThe fingerprint, automation traces such as navigator.webdriver, headless traits, and whether the user agent, Client Hints and engine features agree
BehaviourHow the visitor actsMouse paths, scrolling, typing rhythm, the order and timing of page views
ChallengesA task that is cheap for a person or a real browser and costly at scaleCAPTCHAs, proof-of-work puzzles, JavaScript challenges that must run in a real browser

The layers back each other up. A request that looks fine on one layer can still give itself away on another, which is why a single fix rarely holds. The breakdown by layer is in how anti-bot systems detect automation.

Common bot detection techniques

  • IP and network reputation. Addresses from hosting providers, known proxies and previously abusive ranges start with a higher risk score.
  • Rate limits and velocity. Too many requests, too regular, or too many accounts from one source.
  • TLS and HTTP/2 fingerprints. The handshake shows which software opened the connection, before any page loads (JA3/JA4, HTTP/2).
  • Header checks. Missing headers, an unusual order, or a user agent that the Client Hints contradict.
  • JavaScript fingerprinting. Canvas, WebGL, audio, fonts and dozens more values, read and compared (browser fingerprinting).
  • Automation traces. navigator.webdriver, side effects of the DevTools Protocol, and functions that have been overridden (navigator.webdriver, CDP detection).
  • Headless traits. Signs that the browser has no real window or graphics (headless detection).
  • Behavioural analysis. Pointer movement, timing and navigation that no person produces (behavioural detection).
  • Honeypots. Links and form fields hidden from people, which only a script follows or fills in.
  • Challenges. CAPTCHAs and proof-of-work, usually shown only when the other signals already look doubtful.

Modern systems feed all of this into one model and decide on the combination, so the same request can be allowed on one page and challenged on a more sensitive one.

Why consistency matters more than any single signal

Detection has moved from looking for one tell to checking whether everything agrees. A browser that claims to be Chrome on Windows has to send Chrome's TLS handshake, Chrome's Client Hints for that version, fonts a Windows machine has, and a GPU whose name matches what it draws, and its workers have to report the same values as the page. Each contradiction raises the score.

That is also why patching values one at a time from JavaScript tends to make a browser more conspicuous rather than less. Clearcote's approach is to set the identity inside the browser engine, so the layers agree with each other. The research article How automation gets caught, layer by layer walks through it, and how detection works explains the design.

Bot detection test: check your own setup

Before you automate anything, check what your browser reports:

Running automation responsibly

If you run crawlers or agents, the fewer problems you cause, the fewer you meet. Identify a crawler where the site expects it, honour robots.txt and rate limits, cache what you have already fetched, avoid collecting personal data you have no basis for, and read the site's terms. When you test your own site, allowlist your automation instead of working around your own protection.

Clearcote is built for privacy, testing, research and lawful automation. It keeps a browser's identity consistent; it does not promise that any particular protection will let it through.

FAQ

Is bot detection the same as a CAPTCHA?

No. A CAPTCHA is one tool a bot detection system can use, usually only when other signals already look doubtful. Most bot detection happens without the visitor seeing anything.

Can bot detection block real people?

Yes. Unusual browsers, privacy tools, shared or VPN addresses and assistive software can all raise a risk score. That is why most systems challenge before they block, and why sites tune them per page.

Does a VPN or a proxy get around bot detection?

It changes only the network layer: the IP address and its reputation. The browser fingerprint, the TLS handshake and the behaviour stay the same, and a mismatch such as a time zone that does not fit the IP can add risk rather than remove it.

How do websites tell good bots from bad ones?

Good bots usually identify themselves in their user agent and can be verified, for example by a reverse DNS lookup on the crawler's IP address or by comparing it with the ranges the operator publishes. Unidentified automation is judged by its signals and behaviour instead.

Is web scraping illegal?

Not as such. Whether a particular scrape is lawful depends on the country, the data (personal data and copyrighted content carry extra rules), how it is accessed, and the site's terms. Bot detection is a site's technical choice, not a legal ruling. This is not legal advice.

How can I test whether my browser looks like a bot?

Run the browser fingerprint test in it. It reports which signals contradict each other and what a detector reads from each, and the JA4 checker covers the network handshake.

Clearcote puts this into practice

An open-source Chromium with fingerprint control compiled into the engine. A drop-in for Playwright & Puppeteer.

Free for one browser with GitHub. No card.