What is bot detection?
Bot detection is how a website decides whether a visitor is a person or an automated program. It reads signals on four layers: the network the request comes from, the browser that sends it, how the visitor behaves, and, when in doubt, a challenge such as a CAPTCHA. The result is usually a risk score, and the site decides per page whether to allow, slow down, challenge or block.
Sites use it to keep out credential stuffing, fake accounts, scalping, ad fraud and heavy scraping, while still letting in the bots they want, such as search engine crawlers, which can be verified. Most of it happens without the visitor noticing.
Modern detection is less about any single tell than about consistency: whether the TLS handshake, the user agent, the Client Hints, the fingerprint and the behaviour all describe one real browser. See how anti-bot systems detect automation and how detection works.
Updated
Good bots, bad bots, and the space in between
Not all automated traffic is unwanted. Search engine crawlers, uptime monitors, feed readers and link previews are bots that sites welcome. Credential stuffing, fake sign-ups, ticket and stock hoarding, ad fraud and aggressive scraping are bots that sites try to stop. A lot of automation sits in between: price monitoring, research crawlers, AI agents acting for a user, and testing tools.
Well-behaved crawlers identify themselves. A site can check a crawler's claim instead of trusting its user agent: Google, for example, says to run a reverse DNS lookup on the visiting IP address, confirm the name ends in googlebot.com, google.com or googleusercontent.com, and look the name up again to make sure it resolves to the same address, or to compare the address with the IP ranges Google publishes. The user agents of common crawlers and HTTP clients are listed with our user agent checker.
So bot detection is rarely a yes-or-no gate. Most systems produce a risk score and let the site decide per page: let the request through, slow it down, ask for a challenge, or block it.
How bot detection works: four layers
| Layer | What is checked | Examples |
|---|---|---|
| Network | Where the request comes from and how the connection was made | IP reputation and hosting ranges, request rate, the TLS and HTTP/2 fingerprint compared with the browser the request claims to be |
| Browser | What the client is, read by JavaScript in the page | The fingerprint, automation traces such as navigator.webdriver, headless traits, and whether the user agent, Client Hints and engine features agree |
| Behaviour | How the visitor acts | Mouse paths, scrolling, typing rhythm, the order and timing of page views |
| Challenges | A task that is cheap for a person or a real browser and costly at scale | CAPTCHAs, proof-of-work puzzles, JavaScript challenges that must run in a real browser |
The layers back each other up. A request that looks fine on one layer can still give itself away on another, which is why a single fix rarely holds. The breakdown by layer is in how anti-bot systems detect automation.
Common bot detection techniques
- IP and network reputation. Addresses from hosting providers, known proxies and previously abusive ranges start with a higher risk score.
- Rate limits and velocity. Too many requests, too regular, or too many accounts from one source.
- TLS and HTTP/2 fingerprints. The handshake shows which software opened the connection, before any page loads (JA3/JA4, HTTP/2).
- Header checks. Missing headers, an unusual order, or a user agent that the Client Hints contradict.
- JavaScript fingerprinting. Canvas, WebGL, audio, fonts and dozens more values, read and compared (browser fingerprinting).
- Automation traces.
navigator.webdriver, side effects of the DevTools Protocol, and functions that have been overridden (navigator.webdriver, CDP detection). - Headless traits. Signs that the browser has no real window or graphics (headless detection).
- Behavioural analysis. Pointer movement, timing and navigation that no person produces (behavioural detection).
- Honeypots. Links and form fields hidden from people, which only a script follows or fills in.
- Challenges. CAPTCHAs and proof-of-work, usually shown only when the other signals already look doubtful.
Modern systems feed all of this into one model and decide on the combination, so the same request can be allowed on one page and challenged on a more sensitive one.
Why consistency matters more than any single signal
Detection has moved from looking for one tell to checking whether everything agrees. A browser that claims to be Chrome on Windows has to send Chrome's TLS handshake, Chrome's Client Hints for that version, fonts a Windows machine has, and a GPU whose name matches what it draws, and its workers have to report the same values as the page. Each contradiction raises the score.
That is also why patching values one at a time from JavaScript tends to make a browser more conspicuous rather than less. Clearcote's approach is to set the identity inside the browser engine, so the layers agree with each other. The research article How automation gets caught, layer by layer walks through it, and how detection works explains the design.
Bot detection test: check your own setup
Before you automate anything, check what your browser reports:
- The browser fingerprint test runs a few hundred coherence checks and explains, for every failure, what a detector concludes from it and how to fix it.
- The JA4 fingerprint checker compares your TLS and HTTP/2 handshake with Chrome's and with common HTTP libraries'.
- The user agent checker compares the User-Agent header, the Client Hints and what scripts and workers read.
- The WebRTC leak test shows whether WebRTC reveals an address other than the one you browse from.
- Guides to the public test sites, such as CreepJS and BrowserScan, and what each one measures.
Running automation responsibly
If you run crawlers or agents, the fewer problems you cause, the fewer you meet. Identify a crawler where the site expects it, honour robots.txt and rate limits, cache what you have already fetched, avoid collecting personal data you have no basis for, and read the site's terms. When you test your own site, allowlist your automation instead of working around your own protection.
Clearcote is built for privacy, testing, research and lawful automation. It keeps a browser's identity consistent; it does not promise that any particular protection will let it through.
