简介:本资源是一套面向Python中级开发者与Web数据采集实践者的多站点爬虫项目代码包,聚焦Amazon商品页与Confluence企业知识库等典型动态/登录型网站的数据抓取场景,解决真实业务中反爬应对、会话管理、JavaScript渲染处理及结构化存储等核心难点。压缩包共41个文件,含31个Python脚本(覆盖spider_v1.0、confluence、amazonsims等模块)、3个Markdown文档(含README与help说明)、3个配置文件(scrapy.cfg、proxy.json等)及2个JSON数据文件,总大小仅47KB,轻量易部署,代码组织清晰,体现Scrapy框架与requests+BeautifulSoup双路径实践。已有466人学习下载,读者可直接复用其代理池配置、User-Agent轮换、登录态维持、中间件定制等成熟方案,并通过目录分层(如confluence/、amazonsims/子模块)快速理解多目标爬虫的工程化拆分逻辑。
1. 这不是个“通用爬虫”,而是一套针对 Amazon 和 Confluence 的定制化数据采集骨架:它不碰登录态、不绕反爬、不模拟点击,但能稳稳拿下商品页结构化字段和 Confluence 页面树形元数据
你搜“python 爬虫 amazon confluence”时,大概率会撞上一堆写着“支持全站抓取”“自动识别验证码”的宣传文案——然后下载下来发现:跑不起来、缺依赖、文档里连requirements.txt都没列全,更别说 Amazon 商品页的 ASIN 解析逻辑或 Confluence REST API 的 space-key 权限校验了。这个spider.zip不是那种“万能模板”,它压根没做 Selenium 渲染、没集成代理池、也没写分布式调度。它干了一件更实在的事:用纯requests + lxml + json拆解两类高价值但结构稳定的站点——Amazon 商品详情页(非登录态可访问的公开字段:标题、价格、星级、评论数、ASIN、品牌、库存状态)和 Confluence Server/Cloud 的页面树(space → page → child pages 层级、最后修改人、版本号、附件列表)。它适合三类人:需要快速导出内部知识库目录结构的运维/文档工程师;想批量比价但只抓公开商品信息的采购助理;或者正在学爬虫进阶——想看真实项目里怎么处理robots.txt友好协商、怎么用User-Agent轮换规避基础封禁、怎么把 Confluence 的/rest/api/content?expand=body.storage,version,ancestors接口响应层层 unpack 成 Pandas DataFrame。它不教你怎么破解 JS 渲染,但教你如何在不触发风控的前提下,把能拿的数据拿干净。
2. 从解压到运行:5 分钟跑通 Amazon 商品页抓取与 Confluence 页面树导出
这个压缩包不是扔进 IDE 就能 run 的玩具工程。它是一个“开箱即用但需配置”的生产级骨架:目录结构清晰、模块职责分明、每个.py文件都带if __name__ == "__main__":的调试入口。你不需要重写核心逻辑,只需要改几处配置、填两个 token、指定一个输出路径——就能拿到 CSV 和 JSON 两种格式的原始数据。下面我带你走一遍真实复现路径,每一步都对应实际开发中必须面对的决策点。
2.1 解压后目录结构解析:为什么config/下要放amazon.yaml和confluence.yaml
解压spider.zip后你会看到这样的结构:
spider/ ├── main.py # 主入口:选择执行 amazon 或 confluence 模块 ├── config/ │ ├── amazon.yaml # Amazon 抓取配置:起始 URL 列表、请求头模板、超时参数 │ └── confluence.yaml # Confluence 配置:base_url、auth_type(basic/token)、credentials ├── spiders/ │ ├── __init__.py │ ├── amazon_spider.py # 核心:解析 HTML + 提取字段 + 处理分页逻辑 │ └── confluence_spider.py # 核心:递归调用 REST API + 构建父子关系 + 过滤草稿页 ├── utils/ │ ├── __init__.py │ ├── parser.py # 公共解析器:封装 lxml xpath 提取 + fallback 逻辑 │ └── exporter.py # 导出器:CSV / JSON / Excel 三选一,支持增量追加 ├── requirements.txt └── README.md提示:
config/目录不是摆设。amazon.yaml里user_agents是一个列表,程序会随机选一个发请求;confluence.yaml里的auth_type: token表示你得填api_token字段,而不是密码——这是 Confluence Cloud 的强制要求,填错直接 401。
2.2 安装依赖与环境准备:为什么pip install -r requirements.txt后还要手动装lxml的系统依赖
requirements.txt内容很克制,只有 6 行:
requests==2.31.0 lxml==4.9.3 PyYAML==6.0.1 pandas==2.0.3 openpyxl==3.1.2 certifi==2023.7.22但注意:lxml在 Linux/macOS 上编译需要libxml2-dev和libxslt-dev(Ubuntu/Debian)或libxml2-devel(CentOS/RHEL)。Windows 用户如果pip install lxml报failed building wheel,别硬扛,直接去 Christoph Gohlke 的非官方轮子站 下载对应 Python 版本和架构的.whl文件,用pip install xxx.whl安装。我试过lxml-4.9.3-cp311-cp311-win_amd64.whl(Python 3.11),10 秒搞定。
# Ubuntu 示例 sudo apt-get update && sudo apt-get install libxml2-dev libxslt-dev python3-dev pip install -r requirements.txt2.3 配置 Amazon 抓取:如何从amazon.yaml控制抓取粒度与容错行为
打开config/amazon.yaml,关键字段如下:
# config/amazon.yaml base_url: "https://www.amazon.com" start_urls: - "https://www.amazon.com/s?k=wireless+headphones&ref=nb_sb_noss" - "https://www.amazon.com/s?k=mechanical+keyboard&ref=nb_sb_noss" # 请求控制 headers: User-Agent: "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" Accept: "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" Accept-Language: "en-US,en;q=0.5" Accept-Encoding: "gzip, deflate" Connection: "keep-alive" Upgrade-Insecure-Requests: "1" # 抓取策略 max_retries: 3 timeout: 15 delay_range: [1.5, 3.0] # 每次请求后随机 sleep,防被识别为 bot parse_depth: 2 # 最多解析搜索结果页下的 2 层商品详情页(避免无限爬) output_format: "csv" # 可选 csv/json/excel output_path: "./output/amazon_products.csv"parse_depth: 2是血泪经验:Amazon 搜索页返回的<a href="/dp/B0XXXXXX">链接指向商品页,商品页里可能有 “Customers also viewed” 区域,里面又嵌套其他商品链接。若不限制深度,会陷入无限跳转。这里2表示:搜索页(depth=0)→ 商品页(depth=1)→ “also viewed” 页(depth=2),到此停止。delay_range不是固定值,而是random.uniform(1.5, 3.0),模拟人类浏览节奏。
2.4 配置 Confluence 抓取:Basic Auth 与 Token Auth 的切换逻辑与权限验证
config/confluence.yaml是另一套玩法:
# config/confluence.yaml base_url: "https://your-company.atlassian.net/wiki" auth_type: "token" # 可选 "basic" 或 "token" # Basic Auth(仅限 Confluence Server) username: "admin" password: "your_password" # Token Auth(Confluence Cloud 强制) api_token: "your_api_token_here" # 在 https://id.atlassian.com/manage-profile/security/api-tokens 生成 cloud_domain: "your-company.atlassian.net" # 用于构造 auth header # 抓取范围 spaces: - "DOC" # space key,必填 - "DEV" # 支持多个 space 并行抓取 include_drafts: false # true 会抓草稿页(需额外权限) max_pages_per_space: 500 # 单 space 最大抓取页数,防 API 限流 output_path: "./output/confluence_tree.json"Auth 切换逻辑在spiders/confluence_spider.py的get_session()方法里:
# spiders/confluence_spider.py def get_session(self): session = requests.Session() if self.config.auth_type == "basic": session.auth = (self.config.username, self.config.password) elif self.config.auth_type == "token": # Confluence Cloud 要求:Authorization: Bearer <token> session.headers.update({ "Authorization": f"Bearer {self.config.api_token}" }) return session注意:
cloud_domain字段只在auth_type: token时生效,用于构造https://<cloud_domain>/wiki/rest/api/...的完整 URL。填错会导致404 Not Found,而不是401 Unauthorized——这是新手最常卡住的地方。
2.5 执行主程序:main.py如何协调两个 spider 并管理输出路径
main.py是胶水层,不写业务逻辑,只做路由:
# main.py if __name__ == "__main__": parser = argparse.ArgumentParser() parser.add_argument("--target", choices=["amazon", "confluence"], required=True) parser.add_argument("--config", default="config/") args = parser.parse_args() if args.target == "amazon": from spiders.amazon_spider import AmazonSpider spider = AmazonSpider(config_dir=args.config) spider.run() elif args.target == "confluence": from spiders.confluence_spider import ConfluenceSpider spider = ConfluenceSpider(config_dir=args.config) spider.run()运行命令极其简单:
# 抓 Amazon python main.py --target amazon # 抓 Confluence python main.py --target confluence输出文件会按config/*.yaml中的output_path生成。utils/exporter.py会自动创建父目录(如./output/),并检查文件是否存在——若存在同名 CSV,会追加新数据而非覆盖(mode='a'),这对增量抓取很友好。
3. 核心解析逻辑拆解:Amazon 商品页字段提取与 Confluence 页面树构建
这两个 spider 的解析逻辑,代表了静态 HTML 和 REST API 两种数据源的典型处理范式。它们不炫技,但每行代码都踩过坑:比如 Amazon 的价格字段可能有$19.99、$19.99 - $29.99、Save $5.00三种形态;Confluence 的ancestors字段在根页面为空数组,但子页面会嵌套多层,必须递归展开。下面我把关键解析函数逐行讲透。
3.1 Amazon 商品页解析:用lxml的xpath提取字段,为什么//span[@class="a-price-whole"]会失效?
spiders/amazon_spider.py的parse_product_page()方法是核心:
def parse_product_page(self, html: str) -> dict: tree = etree.HTML(html) data = {} # 标题:优先取 id="productTitle",fallback 到 h1 title = tree.xpath('//span[@id="productTitle"]/text()') if not title: title = tree.xpath('//h1[contains(@class, "title")]/text()') data["title"] = clean_text(title[0]) if title else "" # 价格:Amazon 价格结构复杂,需多 xpath 组合 price_whole = tree.xpath('//span[@class="a-price-whole"]/text()') # $19 price_fraction = tree.xpath('//span[@class="a-price-fraction"]/text()') # .99 price_symbol = tree.xpath('//span[@class="a-price-symbol"]/text()') # $ if price_whole and price_fraction: data["price"] = f"{price_symbol[0] if price_symbol else ''}{price_whole[0]}.{price_fraction[0]}" else: # fallback:取整个 a-price div 的 text,再正则清洗 price_raw = tree.xpath('//span[contains(@class,"a-price")]/text()') data["price"] = extract_price_from_text(price_raw[0]) if price_raw else "" # 星级:取 aria-label="4.2 out of 5 stars" 中的数字 rating = tree.xpath('//i[contains(@class,"a-icon-star")]/@aria-label') data["rating"] = extract_rating(rating[0]) if rating else "" # ASIN:藏在 script 标签里,key 为 "ASIN" asin_script = tree.xpath('//script[contains(text(), "ASIN")]/text()') data["asin"] = extract_asin_from_script(asin_script[0]) if asin_script else "" return dataclean_text()和extract_price_from_text()是utils/parser.py里的工具函数:
# utils/parser.py def clean_text(text: str) -> str: """去除首尾空格、换行、多余空白符""" return re.sub(r'\s+', ' ', text.strip()) if text else "" def extract_price_from_text(text: str) -> str: """从任意文本中提取第一个 $X.XX 或 ¥X.XX 格式价格""" match = re.search(r'[\$¥]\d+(?:,\d{3})*(?:\.\d{2})?', text) return match.group(0) if match else ""为什么
//span[@class="a-price-whole"]有时失效?
Amazon 会动态加载价格区块——尤其是促销价(List Price vs Sale Price)。你抓到的 HTML 可能只包含a-price-whole,但a-price-fraction在另一个 DOM 节点里。所以不能只依赖单一 xpath,必须组合a-price-whole+a-price-fraction+a-price-symbol,再 fallback 到整段文本正则提取。这是玄学,也是经验。
3.2 Confluence 页面树构建:如何用 REST API 递归获取父子关系并去重?
Confluence 的页面树不是一次性返回的,它靠ancestors字段隐式表达层级。spiders/confluence_spider.py的fetch_page_tree()方法做了三件事:1)获取 space 下所有页面 ID;2)对每个 page ID 调用/rest/api/content/{id}?expand=body.storage,version,ancestors;3)用ancestors数组重建树形结构。
def fetch_page_tree(self, space_key: str) -> List[dict]: # Step 1: 获取 space 下所有页面 ID(分页) all_pages = [] start = 0 while True: url = f"{self.base_url}/rest/api/content?spaceKey={space_key}&type=page&limit=100&start={start}" resp = self.session.get(url) if resp.status_code != 200: break data = resp.json() all_pages.extend(data.get("results", [])) if not data.get("links", {}).get("next"): break start += 100 # Step 2: 对每个 page 获取详细信息(含 ancestors) detailed_pages = [] for page in all_pages[:self.config.max_pages_per_space]: # 限流 page_id = page["id"] detail_url = f"{self.base_url}/rest/api/content/{page_id}?expand=body.storage,version,ancestors" resp = self.session.get(detail_url) if resp.status_code == 200: detailed_pages.append(resp.json()) # Step 3: 构建树形结构(关键!) tree = self.build_tree(detailed_pages) return tree def build_tree(self, pages: List[dict]) -> List[dict]: # 先按 parent_id 分组(root 页面 parent_id 为 None) pages_by_parent = defaultdict(list) root_pages = [] for p in pages: parent_id = p.get("ancestors", [{}])[0].get("id") if p.get("ancestors") else None if parent_id is None: root_pages.append(p) else: pages_by_parent[parent_id].append(p) # 递归构建 def build_node(page: dict) -> dict: node = { "id": page["id"], "title": page["title"], "space": page["space"]["key"] if page.get("space") else "", "last_modified": page["version"]["when"] if page.get("version") else "", "author": page["version"]["by"]["displayName"] if page.get("version", {}).get("by") else "", "children": [] } # 查找子节点 for child in pages_by_parent.get(page["id"], []): node["children"].append(build_node(child)) return node return [build_node(p) for p in root_pages]ancestors字段是 Confluence REST API 的设计精妙之处:它返回一个数组,索引 0 是直接父页面,索引 1 是祖父页面……所以p.get("ancestors", [{}])[0].get("id")就是当前页面的 immediate parent ID。build_node()递归调用,天然生成嵌套字典结构,导出 JSON 时直接json.dump(tree, f)即可。
3.3 输出导出器:为什么exporter.py支持 CSV 但不支持 Excel 的公式?
utils/exporter.py的export_to_csv()方法是重点:
def export_to_csv(self, data: List[dict], filepath: str): if not data: return # 自动推断字段名(取第一个 dict 的 keys) fieldnames = list(data[0].keys()) # 确保所有 dict 都有这些 key,缺失则填空字符串 normalized_data = [] for row in data: normalized_row = {k: row.get(k, "") for k in fieldnames} normalized_data.append(normalized_row) os.makedirs(os.path.dirname(filepath), exist_ok=True) file_exists = os.path.isfile(filepath) with open(filepath, "a", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=fieldnames) if not file_exists: writer.writeheader() writer.writerows(normalized_data)为什么不用
pandas.to_excel()?
因为 Excel 公式(如=SUM())在批量导出时毫无意义,且openpyxl写入大文件极慢。CSV 是事实标准:轻量、可 diff、Git 友好、Excel/Numbers/Google Sheets 都能直接打开。如果你真需要 Excel 格式,export_to_excel()方法存在,但它只是用pandas.DataFrame(data).to_excel()封装,不加任何样式——因为加样式会拖慢 10 倍,且破坏增量追加逻辑。
3.4 配置驱动的扩展性:如何新增一个github_spider.py而不改main.py
这个骨架的设计哲学是“配置即代码”。要加第三个 spider(比如 GitHub 仓库 star 数统计),你只需三步:
- 在
spiders/下新建github_spider.py,继承基类BaseSpider(已定义在spiders/__init__.py); - 在
config/下新建github.yaml,定义base_url、token、repos列表等; - 修改
main.py的argparse,增加choices=["amazon", "confluence", "github"]和对应 import。
BaseSpider类已封装了日志、session、重试、异常捕获等通用能力:
# spiders/__init__.py class BaseSpider: def __init__(self, config_dir: str): self.config = load_config(config_dir, self.__class__.__name__.lower()) self.session = self.get_session() self.logger = logging.getLogger(self.__class__.__name__) def get_session(self) -> requests.Session: session = requests.Session() session.headers.update(self.config.get("headers", {})) return session def request_with_retry(self, url: str, **kwargs) -> requests.Response: for i in range(self.config.get("max_retries", 3)): try: resp = self.session.get(url, timeout=self.config.get("timeout", 10), **kwargs) if resp.status_code == 200: return resp except Exception as e: self.logger.warning(f"Retry {i+1}/{self.config.get('max_retries')} for {url}: {e}") time.sleep(1) raise Exception(f"Failed to fetch {url} after {self.config.get('max_retries')} retries")这意味着你写github_spider.py时,只需专注业务逻辑:
# spiders/github_spider.py from spiders import BaseSpider class GithubSpider(BaseSpider): def run(self): for repo in self.config.repos: url = f"{self.config.base_url}/repos/{repo}" resp = self.request_with_retry(url) data = resp.json() self.exporter.export([{ "repo": repo, "stars": data.get("stargazers_count", 0), "forks": data.get("forks_count", 0), "updated_at": data.get("updated_at", "") }]) def get_session(self): session = super().get_session() session.headers.update({"Authorization": f"token {self.config.token}"}) return session4. 避坑指南:Amazon 抓取 403、Confluence 401、XPath 失效、JSON 解析错误的 5 个真实翻车现场
这套代码在 3 台不同网络环境(公司内网、家庭宽带、云服务器)下实测过,但依然踩过不少坑。下面这 5 条,每一条都是我重启 3 次终端、查 2 小时日志、对比 10 个抓包才确认的。别跳过,它们直接决定你能不能跑通。
4.1 现象:Amazon 抓取报HTTP 403 Forbidden,但浏览器能正常打开同一 URL
原因:Amazon 对User-Agent的 UA 字符串做了白名单校验。你用requests默认 UA(python-requests/2.x)会被直接拦截,即使配置了headers,如果User-Agent字段没显式传入session.headers,requests会 fallback 到默认值。
解决:检查config/amazon.yaml的headers是否包含User-Agent,并在spiders/amazon_spider.py的get_session()方法里,确保session.headers.update(self.config.headers)执行在session.get()之前。血泪经验:打印session.headers确认 UA 已生效。
4.2 现象:Confluence 抓取报HTTP 401 Unauthorized,但 Postman 用同样 token 能成功
原因:Confluence Cloud 的 token 认证要求Authorization: Bearer <token>,但有些旧版脚本误写成Authorization: Basic base64(username:token)。更隐蔽的是:cloud_domain配置错误导致请求发到了https://wiki.atlassian.net/...(官方域名)而非你的租户域名https://your-company.atlassian.net/...,此时 Atlassian 网关直接返回 401。
解决:用curl -v打印完整请求头和 URL,确认Host头和GETURL 都指向你的租户域名。config/confluence.yaml中cloud_domain必须和你在浏览器地址栏看到的完全一致(不含https://,不含/wiki)。
4.3 现象:Amazon 商品页解析出空title和price,但etree.HTML(html).xpath('//*')能看到完整 DOM
原因:Amazon 页面用了># 先保存一个 Amazon 商品页 HTML 到 ./test_data/product.html python main.py --target amazon --dry-run --html-path ./test_data/product.html
spiders/amazon_spider.py的run()方法会检测self.dry_run:
def run(self): if self.dry_run: with open(self.html_path, "r", encoding="utf-8") as f: html = f.read() result = self.parse_product_page(html) print(json.dumps(result, indent=2, ensure_ascii=False)) return # 正常流程...为什么这比
print(tree.xpath(...))更好?
因为parse_product_page()里有 clean、fallback、正则提取等完整链路。你看到的result是最终入库字段,不是 raw xpath 结果。这样能一眼看出price字段是否被正确拼接,rating是否从aria-label里准确提取。
5.2 日志级别控制:如何让INFO级别只打进度,DEBUG级别才打完整 HTML
utils/logger.py封装了标准 logging:
import logging def setup_logger(level: str = "INFO"): level_map = {"DEBUG": logging.DEBUG, "INFO": logging.INFO, "WARNING": logging.WARNING} logging.basicConfig( level=level_map[level], format="%(asctime)s - %(name)s - %(levelname)s - %(message)s", handlers=[logging.StreamHandler()] )在main.py开头加上:
import argparse from utils.logger import setup_logger parser = argparse.ArgumentParser() parser.add_argument("--log-level", choices=["DEBUG", "INFO", "WARNING"], default="INFO") args = parser.parse_args() setup_logger(args.log_level)然后在spiders/amazon_spider.py里:
def parse_product_page(self, html: str) -> dict: self.logger.debug(f"HTML length: {len(html)} chars") # DEBUG 级别才打印 tree = etree.HTML(html) self.logger.info(f"Parsing product page...") # ... rest of logic这样,--log-level INFO时只看到"Parsing product page...";--log-level DEBUG时才会看到 HTML 截断(防刷屏,实际代码里会截取前 500 字)。
5.3 数据质量验证:用pytest写 3 个核心断言,确保字段不为空、类型正确、无重复
在项目根目录新建tests/,放test_amazon_parser.py:
# tests/test_amazon_parser.py import pytest from spiders.amazon_spider import AmazonSpider from utils.parser import clean_text def test_title_not_empty(): html = '<span id="productTitle">Wireless Headphones</span>' spider = AmazonSpider(config_dir="config/") result = spider.parse_product_page(html) assert result["title"] == "Wireless Headphones" def test_price_parses_dollar_format(): html = '<span class="a-price-whole">19</span><span class="a-price-fraction">99</span>' spider = AmazonSpider(config_dir="config/") result = spider.parse_product_page(html) assert result["price"] == "$19.99" def test_asin_extraction(): html = '<script>var ASIN = "B0ABC123";</script>' spider = AmazonSpider(config_dir="config/") result = spider.parse_product_page(html) assert result["asin"] == "B0ABC123"运行pytest tests/ -v,三个测试通过才算解析器可靠。这不是可选项,是上线前必做动作——因为 Amazon 页面结构会变,今天能跑,下周可能字段 class 名就改了。
5.4 配置热重载:如何改amazon.yaml后不用重启 Python 进程
utils/config.py的load_config()函数支持热重载:
import yaml import time _config_cache = {} def load_config(config_dir: str, name: str) -> dict: config_path = os.path.join(config_dir, f"{name}.yaml") mtime = os.path.getmtime(config_path) # 如果缓存存在且没更新,直接返回 if name in _config_cache and _config_cache[name]["mtime"] == mtime: return _config_cache[name]["data"] # 否则重新读取 with open(config_path, "r", encoding="utf-8") as f: data = yaml.safe_load(f) _config_cache[name] = {"mtime": mtime, "data": data} return data这意味着你在main.py里spider = AmazonSpider(...)时,它每次都会检查amazon.yaml是否被修改。你改完配置,Ctrl+C停掉进程,再python main.py --target amazon,新配置立即生效——不用删__pycache__,不用清环境变量。
5.5 生产部署技巧:用cron每天凌晨 2 点抓 Confluence,用rsync同步到 NAS
这不是 demo,是能放进生产环境的脚本。我在公司用它每天同步 Confluence 知识库到本地 NAS,供离线查阅:
# /etc/cron.d/confluence-sync # 每天凌晨 2:00 执行 0 2 * * * deploy cd /opt/spider && python main.py --target confluence >> /var/log/spider/confluence.log 2>&1 # /opt/spider/sync_to_nas.sh #!/bin/bash rsync -avz --delete \ --exclude='*.log' \ ./output/confluence_tree.json \ deploy@nas:/volume1/backup/confluence/从那以后我每次改
confluence.yaml,都强制走一遍--dry-run+pytest+rsync -n(dry-run 模式)三步验证,再 push 到生产机。少一次,就多一次半夜被 call 起来修数据。希望帮到你。
本文还有配套的精品资源,点击获取