消息首页搜索举报

Python网络数据采集(第2版第二版版英文版) [美] 瑞安·米切尔东南大学出版社 9787564179779 正版旧书

正版旧书里面部分笔记内容完好可正常使用旧书不附带光盘

15.3 九五品

库存3件

江西南昌

认证卖家担保交易快速发货售后保障

作者[美] 瑞安·米切尔

出版社东南大学出版社

ISBN9787564179779

出版时间2018-11

装帧线装

页数288页

货号4231089

上书时间2024-04-25

辉煌二手教材专营店

已实名已认证进店收藏店铺

在售商品暂无
平均发货时间 10小时
好评率暂无

最新上架

运作何常在北京联合出版公司 9787559614445 正版旧书

运作何常在北京联合出版公司 9787559614445 正版旧书 ¥6.50

智慧港口实践宁涛人民邮电出版社 9787115543271 正版旧书

智慧港口实践宁涛人民邮电出版社 9787115543271 正版旧书 ¥24.20

居里夫人传(新编语文教材指定阅读书系) 玛丽·居里长江文艺出版社 9787535497314 正版旧书

居里夫人传(新编语文教材指定阅读书系) 玛丽·居里长江文艺出版社 9787535497314 正版旧书 ¥3.40

少年特种兵海岛特种战系列(2)-岛屿交锋张永军中国少年儿童出版社 9787514808049 正版旧书

少年特种兵海岛特种战系列(2)-岛屿交锋张永军中国少年儿童出版社 9787514808049 正版旧书 ¥3.40

电脑应用技巧:卓越版本书编委会电子工业出版社 9787121019449 正版旧书

电脑应用技巧:卓越版本书编委会电子工业出版社 9787121019449 正版旧书 ¥2.60

林清玄说禅之三·好雪片片林清玄海南出版社 9787544325950 正版旧书

林清玄说禅之三·好雪片片林清玄海南出版社 9787544325950 正版旧书 ¥5.90

中药安全与合理应用导论张冰中国中医药出版社 9787513242264 正版旧书

中药安全与合理应用导论张冰中国中医药出版社 9787513242264 正版旧书 ¥3.40

语文700题胡媛媛湖北美术出版社 9787539484075 正版旧书

语文700题胡媛媛湖北美术出版社 9787539484075 正版旧书 ¥2.60

工程力学王振发科学出版社 9787030087683 正版旧书

工程力学王振发科学出版社 9787030087683 正版旧书 ¥2.60

商品详情

品相描述：九五品

商品描述: 温馨提示：亲！旧书库存变动比较快，有时难免会有断货的情况，为保证您的利益，拍前请务必联系卖家咨询库存情况！谢谢！

书名：Python网络数据采集(第2版版英文版)
编号：4231089
ISBN：9787564179779[十位:]
作者：[美] 瑞安·米切尔
出版社：东南大学出版社
出版日期：2018年11月
页数：288
定价：89.00 元
参考重量：0.500Kg
-------------------------

新旧程度：6-9成新左右，不影响阅读，详细情况请咨询店主
如图书附带、磁带、学习卡等请咨询店主是否齐全

* 图书目录 *

Preface

Part I. Building Scrapers
1. Your First Web Scraper
Connecting
An Introduction to BeautifulSoup
Installing BeautifulSoup
Running BeautifulSoup
Connecting Reliably and Handling Exceptions
2. Advanced HTML Parsing
You Don't Always Need a Hammer
Another Serving of BeautifulSoup
findo and findallo with BeautifulSoup
Other BeautifulSoup Objects
Navigating Trees
Regular Expressions
Regular Expressions and BeautifulSoup
Accessing Attributes
Lambda Expressions
3. Writing Web Crawlers
Traversing a Single Domain
Crawling an Entire Site
Collecting Data Across an Entire Site
Crawling Across the Internet
4. Web Crawling Models
Planning and Defining Objects
Dealing with Different Website Layouts
Structuring Crawlers
Crawling Sites Through Search
Crawling Sites Through Links
Crawling Multiple Page Types
Thinking About Web Crawler Models
5. Scrapy
Installing Scrapy
Initializing a New Spider
Writing a Simple Scraper
Spidering with Rules
Creating Items
Outputting Items
The Item Pipeline
Logging with Scrapy
More Resources
6. St0ring Data
Media Files
Storing Data to CSV
MySQL
Installing MySQL
Some Basic Commands
Integrating with Python
Database Techniques and Good Practice
"Six Degrees" in MySQL
Email

Part II. Advanced Scraping
7. Reading Documents
Document Encoding
Text
Text Encoding and the Global Internet
CSV
Reading CSV Files
PDF
Microsoft Word and .docx
8. Cleaning Your Dirty Data
Cleaning in Code
Data Normalization
Cleaning After the Fact
OpenRefine
9. Reading and Writing Natural Languages
Summarizing Data
Markov Models
Six Degrees of Wikipedia: Conclusion
Natural Language Toolkit
Installation and Setup
Statistical Analysis with NLTK
Lexicographical Analysis with NLTK
Additional Resources
10. Crawling Through Forms and Logins
Python Requests Library
Submitting a Basic Form
Radio Buttons, Checkboxes, and Other Inputs
Submitting Files and Images
Handling Logins and Cookies
HTTP Basic Access Authentication
Other Form Problems
11. Scraping JavaScript
A Brief Introduction to JavaScript
Common JavaScript Libraries
Ajax and Dynamic HTML
Executing JavaScript in Python with Selenium
Additional Selenium Webdrivers
Handling Redirects
A Final Note on JavaScript
12. Crawling Through APIs
A Brief Introduction to APIs
HTTP Methods and APIs
More About API Responses
Parsing JSON
Undocumented APIs
Finding Undocumented APIs
Documenting Undocumented APIs
Finding and Documenting APIs Automatically
Combining APIs with Other Data Sources
More About APIs
13. Image Processing and Text Recognition
Overview of Libraries
Pillow
Tesseract
NumPy
Processing Well-Formatted Text
Adjusting Images Automatically
Scraping Text from Images on Websites
Reading CAPTCHAs and Training Tesseract
Training Tesseract
Retrieving CAPTCHAs and Submitting Solutions
14. Avoiding Scraping Traps
A Note on Ethics
Looking Like a Human
Adjust Your Headers
Handling Cookies with JavaScript
Timing Is Everything
Common Form Security Features
Hidden Input Field Values
Avoiding Honeypots
The Human Checklist
15. Testing Your Website with Scrapers
An Introduction to Testing
What Are Unit Tests?
Python unittest
Testing Wikipedia
Testing with Selenium
Interacting with the Site
unittest or Selenium?
16. Web Crawling in Parallel
Processes versus Threads
Multithreaded Crawling
Race Conditions and Queues
The threading Module
Multiprocess Crawling
Multiprocess Crawling
Communicating Between Processes
Multiprocess Crawling——Another Approach
17. Scraping Rem0tely
Why Use Remote Servers?
Avoiding IP Address Blocking
Portability and Extensibility
Tor
PySocks
Remote Hosting
Running from a Website-Hosting Account
Running from the Cloud
Additional Resources
18. The Legalities and Ethics of Web Scraping
Trademarks, Copyrights, Patents, Oh My!
Copyright Law
Trespass to Chattels
The Computer Fraud and Abuse Act
robots.txt and Terms of Service
Three Web Scrapers
eBay versus Bidder's Edge and Trespass to Chattels
United States v. Auernheimer and The Computer Fraud and Abuse Act
Field v. Google: Copyright and robots.txt
Moving Forward
Index

— 没有更多了 —

店铺评价

消息首页搜索

以下为对购买帮助不大的评价

此功能需要访问孔网APP才能使用

暂时不用

打开孔网APP