Skip to content

How do you scrape a website with Selenium?

To scrape a website with Selenium, install it with pip install selenium, start a Chrome driver, load the page with driver.get(), wait for the content with WebDriverWait, read it with find_elements(By.CSS_SELECTOR, ...), and call driver.quit() when done. Because Selenium drives a real browser, it sees the page after JavaScript has run.

Use it for pages that need a browser: content rendered by JavaScript, data behind clicks, logins or infinite scroll. For pages whose HTML already contains the data, an HTTP client with an HTML parser is much faster. The guide below walks through a complete scraper, waits, pagination, saving the data, best practices, and how Selenium compares with Playwright.

Out of the box, Selenium-driven Chrome is easy for sites to recognise (navigator.webdriver, ChromeDriver traces, headless traits). The guide covers what the Selenium-side tools change, and what they leave alone.

Updated

When to use Selenium for web scraping

Selenium controls a real browser, so it sees the page after JavaScript has run. That makes it the right tool when the content is rendered in the browser, sits behind clicks, logins or infinite scroll, or needs a real browser session. For pages whose HTML already contains the data, a plain HTTP client with an HTML parser (such as Requests with Beautiful Soup) is many times faster and lighter.

Selenium is also not the only browser option. Playwright and Puppeteer drive the browser over the DevTools Protocol, wait for elements automatically and run several pages in one browser; many scrapers that start on Selenium move to them for speed. The comparison is at the end of this guide.

Set up Selenium

Install the package. Since Selenium 4.6, Selenium Manager downloads a matching driver automatically, so you only need Chrome (or Firefox) installed:

pip install selenium

Scrape a page step by step

The example uses the JavaScript version of quotes.toscrape.com, a practice site built for scraping tutorials. Its HTML contains no quotes at all: they are added by JavaScript after the page loads, which is exactly the case a plain HTTP request cannot handle.

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")  # remove to watch the browser
driver = webdriver.Chrome(options=options)

try:
    driver.get("https://quotes.toscrape.com/js/")
    # Wait until JavaScript has added the quotes, instead of sleeping.
    WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".quote"))
    )
    for quote in driver.find_elements(By.CSS_SELECTOR, ".quote"):
        text = quote.find_element(By.CSS_SELECTOR, ".text").text
        author = quote.find_element(By.CSS_SELECTOR, ".author").text
        print(f"{author}: {text}")
finally:
    driver.quit()

Four parts do the work: driver.get loads the page, WebDriverWait waits for the content, find_elements with a CSS selector finds every quote, and .text reads what a person would see. driver.quit() in a finally block makes sure the browser closes even when something fails.

Wait for content instead of sleeping

The most common mistake in Selenium scrapers is time.sleep(): too short and the element is not there yet, too long and every page wastes seconds. An explicit wait polls until a condition holds and continues the moment it does:

  • presence_of_element_located: the element is in the DOM.
  • visibility_of_element_located: it is in the DOM and visible.
  • element_to_be_clickable: visible and enabled, before a click.
  • text_to_be_present_in_element: a value has loaded.

Avoid mixing implicit waits (driver.implicitly_wait) with explicit ones; the two interact in ways that make timeouts hard to predict.

Pagination, clicks and infinite scroll

To follow a Next link, click it and wait for the old content to go stale before reading the new page. For infinite scroll, scroll to the bottom and wait until more items have loaded:

while True:
    quotes = driver.find_elements(By.CSS_SELECTOR, ".quote")
    # ... read the quotes ...
    next_links = driver.find_elements(By.CSS_SELECTOR, "li.next a")
    if not next_links:
        break
    next_links[0].click()
    WebDriverWait(driver, 10).until(EC.staleness_of(quotes[0]))

Save the data

Collect rows as dictionaries and write them once at the end, with the standard library:

import csv

rows = [{"author": "Albert Einstein", "text": "..."}]  # filled in by the scraper

with open("quotes.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["author", "text"])
    writer.writeheader()
    writer.writerows(rows)

For large pages it is often faster to take driver.page_source once and parse it with lxml or Beautiful Soup than to call find_element hundreds of times, because each call is a round trip to the browser.

Common Selenium scraping errors and how to fix them

ErrorUsual causeFix
NoSuchElementExceptionThe selector is wrong, or the element has not been added yetCheck the selector in the browser's developer tools, and wait for the element instead of finding it at once
TimeoutExceptionThe wait's condition never came trueThe element may be in an iframe, need a scroll or a click first, or not exist at the small headless window size
StaleElementReferenceExceptionThe page re-rendered and replaced the element you heldFind the element again after the update, rather than keeping a reference across page changes
ElementClickInterceptedExceptionSomething covers the element: a cookie banner, a sticky header, an overlayClose or wait out the overlay, or scroll the element into view before clicking
SessionNotCreatedExceptionThe driver and the browser versions do not matchUpgrade Selenium (4.6 or newer) and let Selenium Manager fetch the matching driver

Two structures need an extra step. Content inside an <iframe> is a separate document: call driver.switch_to.frame(...) before finding elements in it and driver.switch_to.default_content() afterwards. Content inside a shadow root (web components) is reached through the host element's shadow_root property, then searched with CSS selectors.

Selenium web scraping best practices

  • Run headless on servers, and set a realistic window size with --window-size=1920,1080; the headless default is small.
  • Reuse one driver for many pages rather than starting a browser per URL.
  • Skip images when you do not need them: the Chrome preference profile.managed_default_content_settings.images set to 2.
  • Catch TimeoutException and NoSuchElementException, log the URL, and move on instead of crashing the whole run.
  • Keep the request rate polite, honour robots.txt and the site's terms, and do not collect personal data you have no basis for.

Why Selenium scrapers get blocked

Out of the box, a Selenium-driven Chrome is easy to recognise. Under automation navigator.webdriver is true, ChromeDriver leaves its own traces in the page, a headless session on a server has a small window and often software-rendered graphics, and the fingerprint is that of a data-center machine. See why navigator.webdriver reveals automation and how headless browsers are detected.

The Selenium-side tools hide parts of this: undetected-chromedriver patches the driver, SeleniumBase's UC and CDP modes change how the browser is driven, and selenium-stealth (unmaintained since 2020) overrides a few values with JavaScript. None of them changes the fingerprint the browser itself reports.

Clearcote changes the fingerprint inside the browser engine and is driven with Playwright or Puppeteer. If you want to keep a Selenium-style framework, SeleniumBase accepts a custom browser through binary_location, so Clearcote can run underneath it. Whatever the tool, scrape responsibly: none of them is a licence to ignore a site's rules.

Selenium vs Playwright for web scraping

SeleniumPlaywright
WaitingExplicit waits you writeAutomatic waiting before most actions
ProtocolWebDriver (and WebDriver BiDi)DevTools Protocol for Chromium, its own for Firefox and WebKit
LanguagesPython, Java, C#, JavaScript, Ruby and moreNode, Python, Java, .NET
Parallel pagesOne driver per browserMany isolated contexts in one browser
Network controlLimited without extra toolsIntercept, block and mock requests built in
Best forExisting Selenium code and teams, the widest language supportNew scrapers that need speed and fewer flaky waits

The same scrape in Playwright for Python:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto("https://quotes.toscrape.com/js/")
    for quote in page.locator(".quote").all():  # locators wait on their own
        print(quote.locator(".author").inner_text(), quote.locator(".text").inner_text())
    browser.close()

FAQ

Is Selenium good for web scraping?

Yes, for pages that need a real browser: JavaScript-rendered content, logins and interaction. It is slower and heavier than plain HTTP requests, so use it only where the page needs it, and consider Playwright for new projects.

Can Selenium scrape JavaScript-rendered websites?

Yes. Selenium drives a real browser, so it reads the page after JavaScript has run. Use an explicit wait for the elements the script adds before you read them.

Is Selenium faster than Beautiful Soup?

No. Beautiful Soup only parses HTML you have already downloaded, which takes milliseconds; Selenium starts and drives a whole browser. A common pattern is to load the page with Selenium and parse driver.page_source with Beautiful Soup.

How do I run Selenium headless?

Add --headless=new to the Chrome options: options.add_argument("--headless=new"). Since Chrome 132, plain --headless also starts the new headless mode.

Why is my Selenium scraper detected?

Usually because of automation traces (navigator.webdriver, ChromeDriver artifacts), headless traits, a data-center IP address, or a request rate no person produces. Check your browser with the browser fingerprint test to see which signals stand out.

Is web scraping with Selenium legal?

The tool does not change the answer. Whether a scrape is lawful depends on the country, the data (personal data and copyrighted content carry extra rules), how you access it, and the site's terms. This is not legal advice.

Related reading

Q&AWhat is a headless browser?A headless browser runs a full browser engine with no visible window — same rendering and JavaScript, driven by code. It's the backbone of automation, and a frequent detection target.Q&AWhy does navigator.webdriver reveal automation?navigator.webdriver is true whenever a browser is launched under automation flags; overriding it from JavaScript is itself detectable, so the fix belongs in the engine.Q&AWhat is a scraping browser?A scraping browser is a real browser used to collect data from JavaScript-heavy sites. It is either a hosted service you connect to over CDP, or one you run yourself.Q&APlaywright vs Puppeteer for stealth: does it matter?Both drive Chromium over CDP, so both leak the same automation artifacts. Stealth comes from the browser binary and launch defaults, not the framework you pick.CompareClearcote vs selenium-stealthselenium-stealth still gets downloaded hundreds of thousands of times a month, but its last release was in 2020. It overrides a few values with JavaScript and asks you to type the GPU and platform in by hand. Clearcote generates the whole identity inside the engine, so the values agree with each other and with the network.CompareClearcote vs SeleniumBaseSeleniumBase is a big, well-maintained Python automation framework, and its UC and CDP modes are the most active stealth option in the Selenium world. They change how the browser is driven, not what it reports: the fingerprint is still your machine's. Clearcote changes it in the engine, and SeleniumBase can run on top of it.CompareClearcote vs undetected-chromedriverundetected-chromedriver is still downloaded about 2 million times a month, but its last release was in February 2024. It hides the Selenium/ChromeDriver seam on a stock Chrome and never touches the fingerprint or the network layer. Clearcote changes both in the engine.

Clearcote puts this into practice

An open-source Chromium with fingerprint control compiled into the engine. A drop-in for Playwright & Puppeteer.

Free for one browser with GitHub. No card.