0% found this document useful (0 votes)
24 views6 pages

Web Scraping with Python: Overview

The document is a comprehensive guide on web scraping using Python, covering topics from building basic scrapers to advanced techniques. It includes detailed instructions on using libraries like BeautifulSoup and Scrapy, as well as handling data storage and cleaning. Additionally, it addresses legal and ethical considerations in web scraping.

Uploaded by

Samy Bouria
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views6 pages

Web Scraping with Python: Overview

The document is a comprehensive guide on web scraping using Python, covering topics from building basic scrapers to advanced techniques. It includes detailed instructions on using libraries like BeautifulSoup and Scrapy, as well as handling data storage and cleaning. Additionally, it addresses legal and ethical considerations in web scraping.

Uploaded by

Samy Bouria
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2n

d
Ed
iti
on
Web Scraping
with Python
COLLECTING MORE DATA FROM THE MODERN WEB

Ryan Mitchell
Table of Contents

Preface. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix

Part I. Building Scrapers


1. Your First Web Scraper. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
Connecting 3
An Introduction to BeautifulSoup 6
Installing BeautifulSoup 6
Running BeautifulSoup 8
Connecting Reliably and Handling Exceptions 10

2. Advanced HTML Parsing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15


You Don’t Always Need a Hammer 15
Another Serving of BeautifulSoup 16
find() and find_all() with BeautifulSoup 18
Other BeautifulSoup Objects 20
Navigating Trees 21
Regular Expressions 25
Regular Expressions and BeautifulSoup 29
Accessing Attributes 30
Lambda Expressions 31

3. Writing Web Crawlers. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33


Traversing a Single Domain 33
Crawling an Entire Site 37
Collecting Data Across an Entire Site 40
Crawling Across the Internet 42

4. Web Crawling Models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49


Planning and Defining Objects 50
Dealing with Different Website Layouts 53

iii
Structuring Crawlers 58
Crawling Sites Through Search 58
Crawling Sites Through Links 61
Crawling Multiple Page Types 64
Thinking About Web Crawler Models 65

5. Scrapy. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 67
Installing Scrapy 67
Initializing a New Spider 68
Writing a Simple Scraper 69
Spidering with Rules 70
Creating Items 74
Outputting Items 76
The Item Pipeline 77
Logging with Scrapy 80
More Resources 80

6. Storing Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83
Media Files 83
Storing Data to CSV 86
MySQL 88
Installing MySQL 89
Some Basic Commands 91
Integrating with Python 94
Database Techniques and Good Practice 97
“Six Degrees” in MySQL 100
Email 103

Part II. Advanced Scraping


7. Reading Documents. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107
Document Encoding 107
Text 108
Text Encoding and the Global Internet 109
CSV 113
Reading CSV Files 113
PDF 115
Microsoft Word and .docx 117

8. Cleaning Your Dirty Data. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121


Cleaning in Code 121

iv | Table of Contents
Data Normalization 124
Cleaning After the Fact 126
OpenRefine 126

9. Reading and Writing Natural Languages. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 131


Summarizing Data 132
Markov Models 135
Six Degrees of Wikipedia: Conclusion 139
Natural Language Toolkit 142
Installation and Setup 142
Statistical Analysis with NLTK 143
Lexicographical Analysis with NLTK 145
Additional Resources 149

10. Crawling Through Forms and Logins. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151


Python Requests Library 151
Submitting a Basic Form 152
Radio Buttons, Checkboxes, and Other Inputs 154
Submitting Files and Images 155
Handling Logins and Cookies 156
HTTP Basic Access Authentication 157
Other Form Problems 158

11. Scraping JavaScript. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161


A Brief Introduction to JavaScript 162
Common JavaScript Libraries 163
Ajax and Dynamic HTML 165
Executing JavaScript in Python with Selenium 166
Additional Selenium Webdrivers 171
Handling Redirects 171
A Final Note on JavaScript 173

12. Crawling Through APIs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 175


A Brief Introduction to APIs 175
HTTP Methods and APIs 177
More About API Responses 178
Parsing JSON 179
Undocumented APIs 181
Finding Undocumented APIs 182
Documenting Undocumented APIs 184
Finding and Documenting APIs Automatically 184
Combining APIs with Other Data Sources 187

Table of Contents | v
More About APIs 190

13. Image Processing and Text Recognition. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 193


Overview of Libraries 194
Pillow 194
Tesseract 195
NumPy 197
Processing Well-Formatted Text 197
Adjusting Images Automatically 200
Scraping Text from Images on Websites 203
Reading CAPTCHAs and Training Tesseract 206
Training Tesseract 207
Retrieving CAPTCHAs and Submitting Solutions 211

14. Avoiding Scraping Traps. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215


A Note on Ethics 215
Looking Like a Human 216
Adjust Your Headers 217
Handling Cookies with JavaScript 218
Timing Is Everything 220
Common Form Security Features 221
Hidden Input Field Values 221
Avoiding Honeypots 223
The Human Checklist 224

15. Testing Your Website with Scrapers. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 227


An Introduction to Testing 227
What Are Unit Tests? 228
Python unittest 228
Testing Wikipedia 230
Testing with Selenium 233
Interacting with the Site 233
unittest or Selenium? 236

16. Web Crawling in Parallel. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 239


Processes versus Threads 239
Multithreaded Crawling 240
Race Conditions and Queues 242
The threading Module 245
Multiprocess Crawling 247
Multiprocess Crawling 249
Communicating Between Processes 251

vi | Table of Contents
Multiprocess Crawling—Another Approach 253

17. Scraping Remotely. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 255


Why Use Remote Servers? 255
Avoiding IP Address Blocking 256
Portability and Extensibility 257
Tor 257
PySocks 259
Remote Hosting 259
Running from a Website-Hosting Account 260
Running from the Cloud 261
Additional Resources 262

18. The Legalities and Ethics of Web Scraping. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 263


Trademarks, Copyrights, Patents, Oh My! 263
Copyright Law 264
Trespass to Chattels 266
The Computer Fraud and Abuse Act 268
[Link] and Terms of Service 269
Three Web Scrapers 272
eBay versus Bidder’s Edge and Trespass to Chattels 272
United States v. Auernheimer and The Computer Fraud and Abuse Act 274
Field v. Google: Copyright and [Link] 275
Moving Forward 276

Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 279

Table of Contents | vii

Common questions

Powered by AI

Scrapy is a robust framework for building web crawlers due to its modularity and efficiency in rule-based scraping . It uses spiders to travel across websites, and its item pipeline feature systematically handles data output by processing items collected in sequential stages, ensuring data transformations and storage in consistent formats, such as CSV or databases, thereby automating complex scraping tasks .

Web crawlers deploy models that include traversing single domains, crawling entire sites, and gathering data across the internet . These models allow adaptation to different layout structures by planning and defining objects that cater to specific site characteristics. Crawlers also vary in strategies, such as through links, searches, or handling multiple page types, thereby optimizing data collection based on site requirements .

API integration provides structured endpoints for accessing data, facilitating targeted data collection without parsing HTML . The use of undocumented APIs poses challenges such as lack of official support or changes without notice, requiring scrapers to reverse-engineer requests and responses, complicating maintenance and increasing the risk of violations of TERMS of SERVICE agreements .

Ethical considerations in web scraping involve respecting copyright laws, TERMS of SERVICE agreements, and the implications of robots.txt files . These legal frameworks protect intellectual properties and site resources, requiring scrapers to act within boundaries of legality, like avoiding undue load on servers and respecting data privacy, thus balancing data collection with ethical responsibilities and avoiding legal actions .

Web scrapers navigate CAPTCHA challenges by training OCR models such as Tesseract for text recognition . To avoid detection by anti-bot systems, scrapers can simulate user behavior through randomized requests, delay tactics, and mimicking browser headers to disguise as human traffic, thereby reducing the likelihood of triggering automated defenses .

Data encoding and document formats, such as CSV or PDF, can lead to errors in data reading due to varying encoding standards and format complexities . Mitigating these challenges involves ensuring proper text and encoding settings, using libraries capable of handling diverse formats, and employing tools like OpenRefine for post-collection cleaning and normalization .

The Python Requests library is essential for interacting with web forms by allowing the submission of data through HTTP requests, which includes handling inputs like radio buttons and checkboxes . Potential issues include handling cookies, authentication requirements, and managing file submissions accurately, which are crucial for maintaining session consistency and fulfilling form constraints effectively .

Multithreaded crawling improves performance by enabling concurrent operations within a single process, reducing time spent on I/O waits . However, it risks race conditions and deadlocks if not managed properly. Multiprocess crawling offers true parallel execution by utilizing multiple cores, aiding in task isolation, but introduces complexity in inter-process communication and higher overhead, impacting overall efficiency if improperly orchestrated .

AJAX signifies dynamic web content that updates asynchronously, posing challenges for static scrapers . Selenium helps by automating browsers to interact with JavaScript-heavy sites, executing scripts to fetch dynamically loaded content, thereby offering a way to scrape sites that rely on client-side rendering and AJAX for submitting user commands and retrieving asynchronous data .

BeautifulSoup offers features such as find() and find_all() methods that allow precise extraction of elements from HTML documents . These methods simplify locating elements by their tag names and attributes, which enhances the accuracy and efficiency of data collection from web pages. Additionally, BeautifulSoup facilitates navigating HTML trees, enabling scrapers to efficiently traverse and access different elements .

You might also like