简介:本资源是一套完整的Python数据科学实战项目,面向数据分析初学者与编程实践者,聚焦美食领域从数据获取到可视化呈现的全流程闭环。项目涵盖网络爬虫(requests+BeautifulSoup)、数据清洗分析(pandas+numpy)及多维可视化(Matplotlib+Seaborn)三大核心能力训练,适用于课程设计、毕业实践或技能进阶学习。压缩包共39个文件,含10个核心Python脚本(如manager.py、models.py、api_1_0模块)、5个HTML/JS前端展示页、5个XML配置与IDE工程文件(.idea/.gitignore等),以及README和数据库迁移相关文件,整体706KB,结构清晰、模块解耦,便于逐层理解工程组织逻辑。已有4087人学习下载,读者可直接运行调试,获得可复用的爬虫模板、标准化数据分析流程、带注释的可视化图表代码及轻量级Flask API接口实现,具备强实操性与教学参考价值。
1. 美食数据闭环实战:从网页抓取到交互式看板,一套代码跑通完整数据链路
你有没有试过:想分析“川菜里哪些食材最常组合出现”,结果卡在第一步——连一份像样的菜品清单都凑不齐?不是缺Python基础,而是缺一个能直接跑起来、带真实数据源、有清洗逻辑、还能一键出图的端到端样板。这个mt_food-master项目就是冲着这个痛点来的:它不是教你怎么写requests.get(),而是把「爬取大众点评/下厨房类站点的菜品页→提取标题/难度/耗时/食材/步骤→存进SQLite→用pandas做频次统计和关联分析→最后用Flask+Plotly搭个可筛选的本地看板」全链路压进一个压缩包里。新手照着manager.py改两行URL就能跑通;熟手能直接拆开models.py和api_1_0/views.py,把数据源换成自己公司的内部菜谱库,或者把static/js/chart.js里的Plotly图表换成ECharts——它不讲原理,只提供可替换、可调试、可验证的生产级脚手架。尤其适合餐饮SaaS产品经理做竞品分析原型、高校食品科学课设、或是想练手但总被“环境配不起来”劝退的转行者。
2. 数据爬取层:绕过反爬、解析动态渲染、结构化存储三步落地
2.1 爬虫核心逻辑:manager.py的调度骨架与utils/scraper.py的实操细节
整个爬取流程由manager.py统一调度,它不直接写HTTP请求,而是调用utils/scraper.py中封装好的FoodScraper类。这种分层设计让后续替换数据源(比如从“下厨房”切到“豆果美食”)只需重写scraper.py里的parse_recipe()方法,而不用动调度逻辑。关键点在于:它默认使用requests-html而非纯requests,因为目标网站大量依赖JavaScript渲染菜品列表——requests拿到的是空壳HTML,而requests-html内置PyQuery+Chromium无头模式,能真实执行JS后抓取最终DOM。
# utils/scraper.py 关键片段 from requests_html import HTMLSession import re class FoodScraper: def __init__(self, base_url="https://www.xiachufang.com"): self.session = HTMLSession() self.base_url = base_url def fetch_page(self, url): # 自动处理重定向和会话保持,比requests更鲁棒 r = self.session.get(url, timeout=15) r.html.render(timeout=20, scrolldown=1) # 渲染滚动加载内容 return r.html def parse_recipe(self, html): # 提取标题:兼容多种HTML结构,用正则兜底 title = html.find('h1.title', first=True) title = title.text.strip() if title else re.search(r'<title>(.*?)</title>', html.html).group(1).strip() # 提取食材列表:定位ul.ingredients-list下的li,过滤空项 ingredients = [] for li in html.find('ul.ingredients-list li'): text = li.text.strip() if text and not re.match(r'^\d+\.?$', text): # 排除序号行 ingredients.append(text) return { 'title': title, 'ingredients': ingredients, 'difficulty': self._extract_difficulty(html), 'cooking_time': self._extract_time(html) }提示:
r.html.render()的scrolldown=1参数是关键——很多美食网站用懒加载,不滚动到底部就拿不到全部菜品。timeout=20是硬性要求,本地Chrome启动慢时容易超时,别盲目缩短。
2.2 反爬对抗策略:User-Agent轮换、请求间隔、Referer伪造
项目没用Scrapy,但scraper.py里埋了轻量级反爬逻辑。它读取config.py中预置的UA池(含Chrome、Firefox、移动端),每次请求随机选一个;同时强制time.sleep(random.uniform(1.2, 2.5)),避免被服务器标记为机器人。更重要的是Referer头——所有请求都带上上一级分类页URL,模拟真实用户点击路径:
# config.py 片段 USER_AGENTS = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36", "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.1 Safari/605.1.15", "Mozilla/5.0 (iPhone; CPU iPhone OS 17_1 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.1 Mobile/15E148 Safari/604.1" ] # scraper.py 中实际调用 headers = { "User-Agent": random.choice(config.USER_AGENTS), "Referer": f"{self.base_url}/category/{category_id}/" # 动态生成Referer } r = self.session.get(url, headers=headers, timeout=15)参数说明:random.uniform(1.2, 2.5)比固定sleep(2)更难被识别;Referer必须和当前请求URL匹配层级,否则部分网站返回403。我一般会先手动访问分类页,复制浏览器Network面板里的真实Referer值,再写进代码——玄学但有效。
2.3 数据落库:Alembic迁移管理 + SQLite轻量存储
爬取的数据不存CSV,而是直写SQLite数据库,由models.py定义ORM模型,migrations/目录用Alembic管理表结构变更。这样做的好处是:后续数据分析能直接用SQL聚合,且支持外键约束(比如菜品和食材的多对多关系)。初始化数据库只需一行命令:
# 在项目根目录执行 python -m alembic revision --autogenerate -m "init tables" python -m alembic upgrade headmodels.py中的关键定义:
# models.py from sqlalchemy import Column, Integer, String, Text, ForeignKey from sqlalchemy.ext.declarative import declarative_base from sqlalchemy.orm import relationship Base = declarative_base() class Recipe(Base): __tablename__ = 'recipes' id = Column(Integer, primary_key=True) title = Column(String(200), nullable=False) difficulty = Column(String(20)) cooking_time = Column(String(50)) # 关联食材:通过中间表recipe_ingredients ingredients = relationship("Ingredient", secondary="recipe_ingredients") class Ingredient(Base): __tablename__ = 'ingredients' id = Column(Integer, primary_key=True) name = Column(String(100), unique=True, index=True) # 加索引加速查询 # 中间表:解决多对多 recipe_ingredients = Table('recipe_ingredients', Base.metadata, Column('recipe_id', Integer, ForeignKey('recipes.id')), Column('ingredient_id', Integer, ForeignKey('ingredients.id')) )注意:Ingredient.name加了unique=True和index=True,这是血泪经验——爬下来可能有“土豆”“马铃薯”“洋芋”三种写法,去重必须靠数据库层约束,不能只靠pandas.drop_duplicates(),否则关联查询会翻车。
3. 数据分析层:清洗、关联、挖掘三阶跃迁
3.1 清洗逻辑:utils/cleaner.py中的食材标准化字典
爬取的食材名五花八门:“五花肉”“五花腩”“梅花肉”“猪五花”,但分析时得归为同一类。项目用utils/cleaner.py实现基于规则的标准化,核心是维护一个映射字典INGREDIENT_MAPPING:
# utils/cleaner.py INGREDIENT_MAPPING = { "五花肉": ["五花腩", "梅花肉", "猪五花", "五花"], "土豆": ["马铃薯", "洋芋", "土豆儿"], "西红柿": ["番茄", "tomato", "西红柿儿"], "青椒": ["甜椒", "彩椒", "灯笼椒"] } def standardize_ingredient(name): name = re.sub(r'[^\w\u4e00-\u9fff]+', '', name) # 去标点空格 for std_name, variants in INGREDIENT_MAPPING.items(): if name in variants or std_name in name or name in std_name: return std_name return name # 未匹配则原样返回清洗入口在manager.py的run_analysis()函数中调用:
# manager.py 片段 from utils.cleaner import standardize_ingredient def run_analysis(): # 从数据库读取原始数据 recipes = session.query(Recipe).all() all_ingredients = [] for recipe in recipes: for raw_ing in recipe.ingredients: std_ing = standardize_ingredient(raw_ing.name) all_ingredients.append(std_ing) # 统计频次 ing_counts = pd.Series(all_ingredients).value_counts() print(ing_counts.head(10)) # 输出前10高频食材参数说明:re.sub(r'[^\w\u4e00-\u9fff]+', '', name)是中文清洗关键——\w匹配英文字母数字下划线,\u4e00-\u9fff是Unicode中文范围,合起来保留中英文和数字,删掉括号、单位(“g”“克”)、描述词(“新鲜”“切丁”)。这步不做,后续热力图会全是噪音。
3.2 关联分析:用Apriori算法挖掘食材共现规律
项目没止步于频次统计,还实现了食材组合挖掘。analysis/association.py用mlxtend库跑Apriori,找“经常一起出现”的食材对:
# analysis/association.py from mlxtend.frequent_patterns import apriori, association_rules import pandas as pd def find_ingredient_pairs(recipes_df): # 构建事务矩阵:每行是一个菜品,每列是一个食材,值为1表示存在 basket = recipes_df['ingredients'].str.join('|').str.get_dummies('|') # Apriori挖掘频繁项集(最小支持度0.02,即2%的菜品包含该组合) frequent_itemsets = apriori(basket, min_support=0.02, use_colnames=True) # 生成关联规则(置信度>0.6) rules = association_rules(frequent_itemsets, metric="confidence", min_threshold=0.6) return rules.sort_values('lift', ascending=False).head(20) # 输出示例:antecedents->consequents, support, confidence, lift # (['鸡蛋'], ['葱']) -> 0.15, 0.82, 3.1注意:
min_support=0.02是经验值。支持度过高(如0.1)只能挖出“盐+油”这种废话组合;过低(如0.005)会产生海量弱规则。我一般先用frequent_itemsets.support.describe()看分布,取25%分位数作为初始值。
3.3 地域风味聚类:TF-IDF + KMeans定位菜系特征
项目还隐藏了一个彩蛋:用TF-IDF向量化菜品标题和食材,再用KMeans聚类,自动发现“川湘辣味”“粤式清淡”“江浙甜鲜”等隐性菜系。代码在analysis/clustering.py:
# analysis/clustering.py from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.cluster import KMeans import jieba def cluster_cuisines(recipes_df): # 中文分词:用jieba切标题+食材列表 def tokenize(text): return ' '.join(jieba.cut(text)) # 合并标题和食材为文本特征 recipes_df['text'] = recipes_df['title'] + ' ' + recipes_df['ingredients'].apply(lambda x: ' '.join(x)) recipes_df['text'] = recipes_df['text'].apply(tokenize) # TF-IDF向量化 vectorizer = TfidfVectorizer(max_features=1000, ngram_range=(1,2)) tfidf_matrix = vectorizer.fit_transform(recipes_df['text']) # KMeans聚类(k=8,对应八大菜系) kmeans = KMeans(n_clusters=8, random_state=42) recipes_df['cluster'] = kmeans.fit_predict(tfidf_matrix) # 输出每个簇的关键词 feature_names = vectorizer.get_feature_names_out() for i in range(8): cluster_terms = tfidf_matrix[recipes_df['cluster']==i].sum(axis=0).A1 top_idx = cluster_terms.argsort()[-10:][::-1] print(f"Cluster {i}: {', '.join([feature_names[j] for j in top_idx])}")参数说明:ngram_range=(1,2)让模型捕捉“豆瓣酱”“花椒油”这类双字词;max_features=1000控制维度,避免稀疏矩阵爆炸;random_state=42保证结果可复现。聚类结果存入数据库,供可视化层按簇筛选。
4. 数据可视化层:Flask后端 + Plotly前端的轻量级看板
4.1 Flask API设计:RESTful接口暴露分析结果
api_1_0/views.py定义了三个核心接口,全部返回JSON,前端直接消费:
GET /api/ingredients/top10:返回高频食材TOP10({'name': '鸡蛋', 'count': 1247})GET /api/associations/rules:返回Apriori规则({'antecedents': ['辣椒'], 'consequents': ['花椒'], 'confidence': 0.78})GET /api/clusters/list:返回聚类结果及各簇代表菜品
# api_1_0/views.py from flask import Blueprint, jsonify from analysis.association import find_ingredient_pairs from analysis.clustering import cluster_cuisines api = Blueprint('api', __name__) @api.route('/ingredients/top10') def top_ingredients(): # 从数据库查,非实时计算,提升响应速度 results = session.execute(""" SELECT i.name, COUNT(*) as cnt FROM ingredients i JOIN recipe_ingredients ri ON i.id = ri.ingredient_id GROUP BY i.name ORDER BY cnt DESC LIMIT 10 """).fetchall() return jsonify([{'name': r[0], 'count': r[1]} for r in results])提示:这里用原生SQL而非ORM,因为聚合查询在SQLite上更快;
LIMIT 10防止前端渲染卡顿。所有接口加了@cache.cached(timeout=300)(需配置Flask-Caching),避免重复计算。
4.2 Plotly图表集成:templates/index.html中的动态渲染
前端用纯HTML+JS,static/js/main.js加载Plotly,调用API绘图:
<!-- templates/index.html --> <div id="top10-chart" style="width: 800px; height: 400px;"></div> <script src="https://cdn.plot.ly/plotly-latest.min.js"></script> <script> fetch('/api/ingredients/top10') .then(r => r.json()) .then(data => { const names = data.map(d => d.name); const counts = data.map(d => d.count); const trace = { x: names, y: counts, type: 'bar', marker: {color: '#FF6B6B'} }; Plotly.newPlot('top10-chart', [trace], { title: '高频食材TOP10', xaxis: {title: '食材'}, yaxis: {title: '出现次数'} }); }); </script>参数说明:Plotly.newPlot()的第三个参数是布局对象,title和轴标签必须显式声明,否则默认为空;marker.color用十六进制色值,避免CSS冲突。所有图表都加了responsive: true(代码中省略),适配不同屏幕。
4.3 交互式筛选:用URL参数驱动后端查询
看板支持按菜系簇筛选,URL形如/dashboard?cluster=3。templates/dashboard.html中的JS监听URL变化,重新请求API:
// static/js/dashboard.js function loadByCluster(clusterId) { fetch(`/api/clusters/recipes?cluster=${clusterId}`) .then(r => r.json()) .then(data => renderRecipeList(data)); } // 页面加载时读取URL参数 const urlParams = new URLSearchParams(window.location.search); const cluster = urlParams.get('cluster') || 'all'; if (cluster !== 'all') { loadByCluster(cluster); }后端api_1_0/views.py对应接口:
@api.route('/clusters/recipes') def recipes_by_cluster(): cluster_id = request.args.get('cluster', type=int) if cluster_id is None: return jsonify([]) recipes = session.query(Recipe).filter(Recipe.cluster == cluster_id).limit(20).all() return jsonify([{ 'title': r.title, 'ingredients': [i.name for i in r.ingredients], 'difficulty': r.difficulty } for r in recipes])注意:
limit(20)是硬性保护,防止一次拉取过多数据拖垮前端。真实项目中应加分页,但本项目为简化,用limit兜底。
5. 避坑指南:爬取失败、数据错乱、图表不显示的五个真实翻车现场
5.1 现象:requests-html渲染超时,报TimeoutError: Waiting for page to load
原因:r.html.render(timeout=20)的20秒不够,尤其网络差或目标站JS复杂时。更隐蔽的是Chromium进程卡死,后续请求全阻塞。
解决:在scraper.py的fetch_page()方法里加进程级超时控制,并捕获异常后重启session:
import signal from contextlib import contextmanager @contextmanager def timeout(seconds): def timeout_handler(signum, frame): raise TimeoutError("Page render timed out") signal.signal(signal.SIGALRM, timeout_handler) signal.alarm(seconds) try: yield finally: signal.alarm(0) def fetch_page(self, url): try: with timeout(30): # 提升到30秒 r = self.session.get(url, timeout=15) r.html.render(timeout=25, scrolldown=1) return r.html except TimeoutError: self.session.close() # 强制关闭旧session self.session = HTMLSession() # 新建session return self.fetch_page(url) # 重试5.2 现象:数据库里食材名重复,recipe_ingredients表出现脏数据
原因:standardize_ingredient()函数没覆盖所有变体,比如“老抽”和“生抽”被当成同一食材,导致关联错误。
解决:在INGREDIENT_MAPPING字典里增加酱油类细分:
"生抽": ["酱油", "浅色酱油", "生抽酱油"], "老抽": ["深色酱油", "老抽酱油", "红酱油"], "蚝油": ["耗油", "耗油汁"]并加校验逻辑:if len(set(ingredients)) < len(ingredients): print("警告:标准化后仍有重复")。
5.3 现象:Apriori结果全是单字词(“盐”“油”“水”),无实际价值
原因:min_support设太高,或事务矩阵构建时没过滤停用词。
解决:在find_ingredient_pairs()前加停用词过滤:
STOPWORDS = {'盐', '油', '水', '料酒', '糖', '鸡精', '味精', '胡椒粉'} basket = basket.drop(columns=[col for col in basket.columns if col in STOPWORDS], errors='ignore')5.4 现象:Flask本地运行正常,部署到Linux服务器后图表空白
原因:Plotly CDN在部分企业内网被拦截,或服务器时间不同步导致HTTPS证书失效。
解决:下载Plotly离线版,放入static/js/目录,改引用:
<!-- 替换CDN链接 --> <script src="{{ url_for('static', filename='js/plotly-2.24.1.min.js') }}"></script>同时检查服务器时间:sudo ntpdate -s time.nist.gov。
5.5 现象:聚类结果每次运行都不一样,无法复现
原因:KMeans随机初始化,random_state没全局统一。
解决:在clustering.py顶部加全局seed,并确保所有随机操作用同一seed:
import numpy as np import random SEED = 42 np.random.seed(SEED) random.seed(SEED) # 所有sklearn模型都传 random_state=SEED6. 进阶技巧:用Docker一键部署看板 + 添加搜索框实时过滤
6.1 Docker化部署:三步打包,彻底解决环境依赖
本地跑通后,用Docker封装成镜像,避免“在我机器上好好的”问题。Dockerfile极简:
FROM python:3.9-slim WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY . . EXPOSE 5000 CMD ["gunicorn", "--bind", "0.0.0.0:5000", "--workers", "2", "app:app"]requirements.txt关键依赖:
Flask==2.3.3 requests-html==0.10.0 pandas==2.1.3 plotly==6.16.0 gunicorn==21.2.0 alembic==1.13.1构建命令:
docker build -t mt-food-dashboard . docker run -p 5000:5000 -v $(pwd)/data:/app/data mt-food-dashboard注意:
-v $(pwd)/data:/app/data将宿主机data/目录挂载到容器内,确保SQLite数据库文件持久化。容器内路径必须和config.py中数据库路径一致(sqlite:///data/app.db)。
6.2 前端增强:添加搜索框,实时过滤食材热力图
templates/dashboard.html增加搜索框和JS逻辑,实现输入即查:
<input type="text" id="search-input" placeholder="搜索食材..." oninput="filterHeatmap(this.value)"> <div id="heatmap" style="width: 900px; height: 500px;"></div> <script> let allData = []; // 全局缓存原始热力图数据 function loadHeatmap() { fetch('/api/associations/heatmap') .then(r => r.json()) .then(data => { allData = data; renderHeatmap(data); }); } function filterHeatmap(keyword) { if (!keyword.trim()) return renderHeatmap(allData); const filtered = allData.filter(item => item.antecedents.includes(keyword) || item.consequents.includes(keyword) ); renderHeatmap(filtered); } function renderHeatmap(data) { // Plotly热力图代码(略,同4.2节结构) } </script>后端新增接口/api/associations/heatmap,返回完整规则列表(非TOP20),供前端自由筛选。
6.3 数据验证技巧:用pytest写三个必跑测试
为防重构破坏核心逻辑,我在tests/目录写了三个轻量测试:
# tests/test_scraper.py def test_ingredient_standardization(): assert standardize_ingredient("五花腩") == "五花肉" assert standardize_ingredient("马铃薯") == "土豆" # tests/test_association.py def test_apriori_min_support(): # 用小样本数据测试,确保支持度阈值生效 sample_basket = pd.DataFrame({ '鸡蛋': [1,1,0,0], '葱': [1,1,1,0], '盐': [1,1,1,1] }) freq = apriori(sample_basket, min_support=0.5, use_colnames=True) assert len(freq) == 2 # 只有鸡蛋&葱、葱&盐满足0.5支持度 # tests/test_api.py def test_api_top10_returns_json(): app.config['TESTING'] = True client = app.test_client() rv = client.get('/api/ingredients/top10') assert rv.status_code == 200 assert isinstance(rv.get_json(), list)运行命令:pytest tests/ -v。从那以后我每次改cleaner.py或association.py,都强制跑这三测——后悔药不如预防针管用。希望帮到你。
本文还有配套的精品资源,点击获取