How do you scrape a website with Selenium?
To scrape a website with Selenium, install it with pip install selenium, start a Chrome driver, load the page with driver.get(), wait for the content with WebDriverWait, read it with find_elements(By.CSS_SELECTOR, ...), and call driver.quit() when done. Because Selenium drives a real browser, it sees the page after JavaScript has run.
Use it for pages that need a browser: content rendered by JavaScript, data behind clicks, logins or infinite scroll. For pages whose HTML already contains the data, an HTTP client with an HTML parser is much faster. The guide below walks through a complete scraper, waits, pagination, saving the data, best practices, and how Selenium compares with Playwright.
Out of the box, Selenium-driven Chrome is easy for sites to recognise (navigator.webdriver, ChromeDriver traces, headless traits). The guide covers what the Selenium-side tools change, and what they leave alone.
Updated
When to use Selenium for web scraping
Selenium controls a real browser, so it sees the page after JavaScript has run. That makes it the right tool when the content is rendered in the browser, sits behind clicks, logins or infinite scroll, or needs a real browser session. For pages whose HTML already contains the data, a plain HTTP client with an HTML parser (such as Requests with Beautiful Soup) is many times faster and lighter.
Selenium is also not the only browser option. Playwright and Puppeteer drive the browser over the DevTools Protocol, wait for elements automatically and run several pages in one browser; many scrapers that start on Selenium move to them for speed. The comparison is at the end of this guide.
Set up Selenium
Install the package. Since Selenium 4.6, Selenium Manager downloads a matching driver automatically, so you only need Chrome (or Firefox) installed:
pip install seleniumScrape a page step by step
The example uses the JavaScript version of quotes.toscrape.com, a practice site built for scraping tutorials. Its HTML contains no quotes at all: they are added by JavaScript after the page loads, which is exactly the case a plain HTTP request cannot handle.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
options = webdriver.ChromeOptions()
options.add_argument("--headless=new") # remove to watch the browser
driver = webdriver.Chrome(options=options)
try:
driver.get("https://quotes.toscrape.com/js/")
# Wait until JavaScript has added the quotes, instead of sleeping.
WebDriverWait(driver, 10).until(
EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".quote"))
)
for quote in driver.find_elements(By.CSS_SELECTOR, ".quote"):
text = quote.find_element(By.CSS_SELECTOR, ".text").text
author = quote.find_element(By.CSS_SELECTOR, ".author").text
print(f"{author}: {text}")
finally:
driver.quit()Four parts do the work: driver.get loads the page, WebDriverWait waits for the content, find_elements with a CSS selector finds every quote, and .text reads what a person would see. driver.quit() in a finally block makes sure the browser closes even when something fails.
Wait for content instead of sleeping
The most common mistake in Selenium scrapers is time.sleep(): too short and the element is not there yet, too long and every page wastes seconds. An explicit wait polls until a condition holds and continues the moment it does:
presence_of_element_located: the element is in the DOM.visibility_of_element_located: it is in the DOM and visible.element_to_be_clickable: visible and enabled, before a click.text_to_be_present_in_element: a value has loaded.
Avoid mixing implicit waits (driver.implicitly_wait) with explicit ones; the two interact in ways that make timeouts hard to predict.
Pagination, clicks and infinite scroll
To follow a Next link, click it and wait for the old content to go stale before reading the new page. For infinite scroll, scroll to the bottom and wait until more items have loaded:
while True:
quotes = driver.find_elements(By.CSS_SELECTOR, ".quote")
# ... read the quotes ...
next_links = driver.find_elements(By.CSS_SELECTOR, "li.next a")
if not next_links:
break
next_links[0].click()
WebDriverWait(driver, 10).until(EC.staleness_of(quotes[0]))Save the data
Collect rows as dictionaries and write them once at the end, with the standard library:
import csv
rows = [{"author": "Albert Einstein", "text": "..."}] # filled in by the scraper
with open("quotes.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["author", "text"])
writer.writeheader()
writer.writerows(rows)For large pages it is often faster to take driver.page_source once and parse it with lxml or Beautiful Soup than to call find_element hundreds of times, because each call is a round trip to the browser.
Common Selenium scraping errors and how to fix them
| Error | Usual cause | Fix |
|---|---|---|
NoSuchElementException | The selector is wrong, or the element has not been added yet | Check the selector in the browser's developer tools, and wait for the element instead of finding it at once |
TimeoutException | The wait's condition never came true | The element may be in an iframe, need a scroll or a click first, or not exist at the small headless window size |
StaleElementReferenceException | The page re-rendered and replaced the element you held | Find the element again after the update, rather than keeping a reference across page changes |
ElementClickInterceptedException | Something covers the element: a cookie banner, a sticky header, an overlay | Close or wait out the overlay, or scroll the element into view before clicking |
SessionNotCreatedException | The driver and the browser versions do not match | Upgrade Selenium (4.6 or newer) and let Selenium Manager fetch the matching driver |
Two structures need an extra step. Content inside an <iframe> is a separate document: call driver.switch_to.frame(...) before finding elements in it and driver.switch_to.default_content() afterwards. Content inside a shadow root (web components) is reached through the host element's shadow_root property, then searched with CSS selectors.
Selenium web scraping best practices
- Run headless on servers, and set a realistic window size with
--window-size=1920,1080; the headless default is small. - Reuse one driver for many pages rather than starting a browser per URL.
- Skip images when you do not need them: the Chrome preference
profile.managed_default_content_settings.imagesset to2. - Catch
TimeoutExceptionandNoSuchElementException, log the URL, and move on instead of crashing the whole run. - Keep the request rate polite, honour robots.txt and the site's terms, and do not collect personal data you have no basis for.
Why Selenium scrapers get blocked
Out of the box, a Selenium-driven Chrome is easy to recognise. Under automation navigator.webdriver is true, ChromeDriver leaves its own traces in the page, a headless session on a server has a small window and often software-rendered graphics, and the fingerprint is that of a data-center machine. See why navigator.webdriver reveals automation and how headless browsers are detected.
The Selenium-side tools hide parts of this: undetected-chromedriver patches the driver, SeleniumBase's UC and CDP modes change how the browser is driven, and selenium-stealth (unmaintained since 2020) overrides a few values with JavaScript. None of them changes the fingerprint the browser itself reports.
Clearcote changes the fingerprint inside the browser engine and is driven with Playwright or Puppeteer. If you want to keep a Selenium-style framework, SeleniumBase accepts a custom browser through binary_location, so Clearcote can run underneath it. Whatever the tool, scrape responsibly: none of them is a licence to ignore a site's rules.
Selenium vs Playwright for web scraping
| Selenium | Playwright | |
|---|---|---|
| Waiting | Explicit waits you write | Automatic waiting before most actions |
| Protocol | WebDriver (and WebDriver BiDi) | DevTools Protocol for Chromium, its own for Firefox and WebKit |
| Languages | Python, Java, C#, JavaScript, Ruby and more | Node, Python, Java, .NET |
| Parallel pages | One driver per browser | Many isolated contexts in one browser |
| Network control | Limited without extra tools | Intercept, block and mock requests built in |
| Best for | Existing Selenium code and teams, the widest language support | New scrapers that need speed and fewer flaky waits |
The same scrape in Playwright for Python:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto("https://quotes.toscrape.com/js/")
for quote in page.locator(".quote").all(): # locators wait on their own
print(quote.locator(".author").inner_text(), quote.locator(".text").inner_text())
browser.close()