简介本资源是一套完整的Python数据科学实战项目面向数据分析初学者与编程实践者聚焦美食领域从数据获取到可视化呈现的全流程闭环。项目涵盖网络爬虫requestsBeautifulSoup、数据清洗分析pandasnumpy及多维可视化MatplotlibSeaborn三大核心能力训练适用于课程设计、毕业实践或技能进阶学习。压缩包共39个文件含10个核心Python脚本如manager.py、models.py、api_1_0模块、5个HTML/JS前端展示页、5个XML配置与IDE工程文件.idea/.gitignore等以及README和数据库迁移相关文件整体706KB结构清晰、模块解耦便于逐层理解工程组织逻辑。已有4087人学习下载读者可直接运行调试获得可复用的爬虫模板、标准化数据分析流程、带注释的可视化图表代码及轻量级Flask API接口实现具备强实操性与教学参考价值。1. 美食数据闭环实战从网页抓取到交互式看板一套代码跑通完整数据链路你有没有试过想分析“川菜里哪些食材最常组合出现”结果卡在第一步——连一份像样的菜品清单都凑不齐不是缺Python基础而是缺一个能直接跑起来、带真实数据源、有清洗逻辑、还能一键出图的端到端样板。这个mt_food-master项目就是冲着这个痛点来的它不是教你怎么写requests.get()而是把「爬取大众点评/下厨房类站点的菜品页→提取标题/难度/耗时/食材/步骤→存进SQLite→用pandas做频次统计和关联分析→最后用FlaskPlotly搭个可筛选的本地看板」全链路压进一个压缩包里。新手照着manager.py改两行URL就能跑通熟手能直接拆开models.py和api_1_0/views.py把数据源换成自己公司的内部菜谱库或者把static/js/chart.js里的Plotly图表换成ECharts——它不讲原理只提供可替换、可调试、可验证的生产级脚手架。尤其适合餐饮SaaS产品经理做竞品分析原型、高校食品科学课设、或是想练手但总被“环境配不起来”劝退的转行者。2. 数据爬取层绕过反爬、解析动态渲染、结构化存储三步落地2.1 爬虫核心逻辑manager.py的调度骨架与utils/scraper.py的实操细节整个爬取流程由manager.py统一调度它不直接写HTTP请求而是调用utils/scraper.py中封装好的FoodScraper类。这种分层设计让后续替换数据源比如从“下厨房”切到“豆果美食”只需重写scraper.py里的parse_recipe()方法而不用动调度逻辑。关键点在于它默认使用requests-html而非纯requests因为目标网站大量依赖JavaScript渲染菜品列表——requests拿到的是空壳HTML而requests-html内置PyQueryChromium无头模式能真实执行JS后抓取最终DOM。# utils/scraper.py 关键片段 from requests_html import HTMLSession import re class FoodScraper: def __init__(self, base_urlhttps://www.xiachufang.com): self.session HTMLSession() self.base_url base_url def fetch_page(self, url): # 自动处理重定向和会话保持比requests更鲁棒 r self.session.get(url, timeout15) r.html.render(timeout20, scrolldown1) # 渲染滚动加载内容 return r.html def parse_recipe(self, html): # 提取标题兼容多种HTML结构用正则兜底 title html.find(h1.title, firstTrue) title title.text.strip() if title else re.search(rtitle(.*?)/title, html.html).group(1).strip() # 提取食材列表定位ul.ingredients-list下的li过滤空项 ingredients [] for li in html.find(ul.ingredients-list li): text li.text.strip() if text and not re.match(r^\d\.?$, text): # 排除序号行 ingredients.append(text) return { title: title, ingredients: ingredients, difficulty: self._extract_difficulty(html), cooking_time: self._extract_time(html) }提示r.html.render()的scrolldown1参数是关键——很多美食网站用懒加载不滚动到底部就拿不到全部菜品。timeout20是硬性要求本地Chrome启动慢时容易超时别盲目缩短。2.2 反爬对抗策略User-Agent轮换、请求间隔、Referer伪造项目没用Scrapy但scraper.py里埋了轻量级反爬逻辑。它读取config.py中预置的UA池含Chrome、Firefox、移动端每次请求随机选一个同时强制time.sleep(random.uniform(1.2, 2.5))避免被服务器标记为机器人。更重要的是Referer头——所有请求都带上上一级分类页URL模拟真实用户点击路径# config.py 片段 USER_AGENTS [ Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36, Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.1 Safari/605.1.15, Mozilla/5.0 (iPhone; CPU iPhone OS 17_1 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.1 Mobile/15E148 Safari/604.1 ] # scraper.py 中实际调用 headers { User-Agent: random.choice(config.USER_AGENTS), Referer: f{self.base_url}/category/{category_id}/ # 动态生成Referer } r self.session.get(url, headersheaders, timeout15)参数说明random.uniform(1.2, 2.5)比固定sleep(2)更难被识别Referer必须和当前请求URL匹配层级否则部分网站返回403。我一般会先手动访问分类页复制浏览器Network面板里的真实Referer值再写进代码——玄学但有效。2.3 数据落库Alembic迁移管理 SQLite轻量存储爬取的数据不存CSV而是直写SQLite数据库由models.py定义ORM模型migrations/目录用Alembic管理表结构变更。这样做的好处是后续数据分析能直接用SQL聚合且支持外键约束比如菜品和食材的多对多关系。初始化数据库只需一行命令# 在项目根目录执行 python -m alembic revision --autogenerate -m init tables python -m alembic upgrade headmodels.py中的关键定义# models.py from sqlalchemy import Column, Integer, String, Text, ForeignKey from sqlalchemy.ext.declarative import declarative_base from sqlalchemy.orm import relationship Base declarative_base() class Recipe(Base): __tablename__ recipes id Column(Integer, primary_keyTrue) title Column(String(200), nullableFalse) difficulty Column(String(20)) cooking_time Column(String(50)) # 关联食材通过中间表recipe_ingredients ingredients relationship(Ingredient, secondaryrecipe_ingredients) class Ingredient(Base): __tablename__ ingredients id Column(Integer, primary_keyTrue) name Column(String(100), uniqueTrue, indexTrue) # 加索引加速查询 # 中间表解决多对多 recipe_ingredients Table(recipe_ingredients, Base.metadata, Column(recipe_id, Integer, ForeignKey(recipes.id)), Column(ingredient_id, Integer, ForeignKey(ingredients.id)) )注意Ingredient.name加了uniqueTrue和indexTrue这是血泪经验——爬下来可能有“土豆”“马铃薯”“洋芋”三种写法去重必须靠数据库层约束不能只靠pandas.drop_duplicates()否则关联查询会翻车。3. 数据分析层清洗、关联、挖掘三阶跃迁3.1 清洗逻辑utils/cleaner.py中的食材标准化字典爬取的食材名五花八门“五花肉”“五花腩”“梅花肉”“猪五花”但分析时得归为同一类。项目用utils/cleaner.py实现基于规则的标准化核心是维护一个映射字典INGREDIENT_MAPPING# utils/cleaner.py INGREDIENT_MAPPING { 五花肉: [五花腩, 梅花肉, 猪五花, 五花], 土豆: [马铃薯, 洋芋, 土豆儿], 西红柿: [番茄, tomato, 西红柿儿], 青椒: [甜椒, 彩椒, 灯笼椒] } def standardize_ingredient(name): name re.sub(r[^\w\u4e00-\u9fff], , name) # 去标点空格 for std_name, variants in INGREDIENT_MAPPING.items(): if name in variants or std_name in name or name in std_name: return std_name return name # 未匹配则原样返回清洗入口在manager.py的run_analysis()函数中调用# manager.py 片段 from utils.cleaner import standardize_ingredient def run_analysis(): # 从数据库读取原始数据 recipes session.query(Recipe).all() all_ingredients [] for recipe in recipes: for raw_ing in recipe.ingredients: std_ing standardize_ingredient(raw_ing.name) all_ingredients.append(std_ing) # 统计频次 ing_counts pd.Series(all_ingredients).value_counts() print(ing_counts.head(10)) # 输出前10高频食材参数说明re.sub(r[^\w\u4e00-\u9fff], , name)是中文清洗关键——\w匹配英文字母数字下划线\u4e00-\u9fff是Unicode中文范围合起来保留中英文和数字删掉括号、单位“g”“克”、描述词“新鲜”“切丁”。这步不做后续热力图会全是噪音。3.2 关联分析用Apriori算法挖掘食材共现规律项目没止步于频次统计还实现了食材组合挖掘。analysis/association.py用mlxtend库跑Apriori找“经常一起出现”的食材对# analysis/association.py from mlxtend.frequent_patterns import apriori, association_rules import pandas as pd def find_ingredient_pairs(recipes_df): # 构建事务矩阵每行是一个菜品每列是一个食材值为1表示存在 basket recipes_df[ingredients].str.join(|).str.get_dummies(|) # Apriori挖掘频繁项集最小支持度0.02即2%的菜品包含该组合 frequent_itemsets apriori(basket, min_support0.02, use_colnamesTrue) # 生成关联规则置信度0.6 rules association_rules(frequent_itemsets, metricconfidence, min_threshold0.6) return rules.sort_values(lift, ascendingFalse).head(20) # 输出示例antecedents-consequents, support, confidence, lift # ([鸡蛋], [葱]) - 0.15, 0.82, 3.1注意min_support0.02是经验值。支持度过高如0.1只能挖出“盐油”这种废话组合过低如0.005会产生海量弱规则。我一般先用frequent_itemsets.support.describe()看分布取25%分位数作为初始值。3.3 地域风味聚类TF-IDF KMeans定位菜系特征项目还隐藏了一个彩蛋用TF-IDF向量化菜品标题和食材再用KMeans聚类自动发现“川湘辣味”“粤式清淡”“江浙甜鲜”等隐性菜系。代码在analysis/clustering.py# analysis/clustering.py from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.cluster import KMeans import jieba def cluster_cuisines(recipes_df): # 中文分词用jieba切标题食材列表 def tokenize(text): return .join(jieba.cut(text)) # 合并标题和食材为文本特征 recipes_df[text] recipes_df[title] recipes_df[ingredients].apply(lambda x: .join(x)) recipes_df[text] recipes_df[text].apply(tokenize) # TF-IDF向量化 vectorizer TfidfVectorizer(max_features1000, ngram_range(1,2)) tfidf_matrix vectorizer.fit_transform(recipes_df[text]) # KMeans聚类k8对应八大菜系 kmeans KMeans(n_clusters8, random_state42) recipes_df[cluster] kmeans.fit_predict(tfidf_matrix) # 输出每个簇的关键词 feature_names vectorizer.get_feature_names_out() for i in range(8): cluster_terms tfidf_matrix[recipes_df[cluster]i].sum(axis0).A1 top_idx cluster_terms.argsort()[-10:][::-1] print(fCluster {i}: {, .join([feature_names[j] for j in top_idx])})参数说明ngram_range(1,2)让模型捕捉“豆瓣酱”“花椒油”这类双字词max_features1000控制维度避免稀疏矩阵爆炸random_state42保证结果可复现。聚类结果存入数据库供可视化层按簇筛选。4. 数据可视化层Flask后端 Plotly前端的轻量级看板4.1 Flask API设计RESTful接口暴露分析结果api_1_0/views.py定义了三个核心接口全部返回JSON前端直接消费GET /api/ingredients/top10返回高频食材TOP10{name: 鸡蛋, count: 1247}GET /api/associations/rules返回Apriori规则{antecedents: [辣椒], consequents: [花椒], confidence: 0.78}GET /api/clusters/list返回聚类结果及各簇代表菜品# api_1_0/views.py from flask import Blueprint, jsonify from analysis.association import find_ingredient_pairs from analysis.clustering import cluster_cuisines api Blueprint(api, __name__) api.route(/ingredients/top10) def top_ingredients(): # 从数据库查非实时计算提升响应速度 results session.execute( SELECT i.name, COUNT(*) as cnt FROM ingredients i JOIN recipe_ingredients ri ON i.id ri.ingredient_id GROUP BY i.name ORDER BY cnt DESC LIMIT 10 ).fetchall() return jsonify([{name: r[0], count: r[1]} for r in results])提示这里用原生SQL而非ORM因为聚合查询在SQLite上更快LIMIT 10防止前端渲染卡顿。所有接口加了cache.cached(timeout300)需配置Flask-Caching避免重复计算。4.2 Plotly图表集成templates/index.html中的动态渲染前端用纯HTMLJSstatic/js/main.js加载Plotly调用API绘图!-- templates/index.html -- div idtop10-chart stylewidth: 800px; height: 400px;/div script srchttps://cdn.plot.ly/plotly-latest.min.js/script script fetch(/api/ingredients/top10) .then(r r.json()) .then(data { const names data.map(d d.name); const counts data.map(d d.count); const trace { x: names, y: counts, type: bar, marker: {color: #FF6B6B} }; Plotly.newPlot(top10-chart, [trace], { title: 高频食材TOP10, xaxis: {title: 食材}, yaxis: {title: 出现次数} }); }); /script参数说明Plotly.newPlot()的第三个参数是布局对象title和轴标签必须显式声明否则默认为空marker.color用十六进制色值避免CSS冲突。所有图表都加了responsive: true代码中省略适配不同屏幕。4.3 交互式筛选用URL参数驱动后端查询看板支持按菜系簇筛选URL形如/dashboard?cluster3。templates/dashboard.html中的JS监听URL变化重新请求API// static/js/dashboard.js function loadByCluster(clusterId) { fetch(/api/clusters/recipes?cluster${clusterId}) .then(r r.json()) .then(data renderRecipeList(data)); } // 页面加载时读取URL参数 const urlParams new URLSearchParams(window.location.search); const cluster urlParams.get(cluster) || all; if (cluster ! all) { loadByCluster(cluster); }后端api_1_0/views.py对应接口api.route(/clusters/recipes) def recipes_by_cluster(): cluster_id request.args.get(cluster, typeint) if cluster_id is None: return jsonify([]) recipes session.query(Recipe).filter(Recipe.cluster cluster_id).limit(20).all() return jsonify([{ title: r.title, ingredients: [i.name for i in r.ingredients], difficulty: r.difficulty } for r in recipes])注意limit(20)是硬性保护防止一次拉取过多数据拖垮前端。真实项目中应加分页但本项目为简化用limit兜底。5. 避坑指南爬取失败、数据错乱、图表不显示的五个真实翻车现场5.1 现象requests-html渲染超时报TimeoutError: Waiting for page to load原因r.html.render(timeout20)的20秒不够尤其网络差或目标站JS复杂时。更隐蔽的是Chromium进程卡死后续请求全阻塞。解决在scraper.py的fetch_page()方法里加进程级超时控制并捕获异常后重启sessionimport signal from contextlib import contextmanager contextmanager def timeout(seconds): def timeout_handler(signum, frame): raise TimeoutError(Page render timed out) signal.signal(signal.SIGALRM, timeout_handler) signal.alarm(seconds) try: yield finally: signal.alarm(0) def fetch_page(self, url): try: with timeout(30): # 提升到30秒 r self.session.get(url, timeout15) r.html.render(timeout25, scrolldown1) return r.html except TimeoutError: self.session.close() # 强制关闭旧session self.session HTMLSession() # 新建session return self.fetch_page(url) # 重试5.2 现象数据库里食材名重复recipe_ingredients表出现脏数据原因standardize_ingredient()函数没覆盖所有变体比如“老抽”和“生抽”被当成同一食材导致关联错误。解决在INGREDIENT_MAPPING字典里增加酱油类细分生抽: [酱油, 浅色酱油, 生抽酱油], 老抽: [深色酱油, 老抽酱油, 红酱油], 蚝油: [耗油, 耗油汁]并加校验逻辑if len(set(ingredients)) len(ingredients): print(警告标准化后仍有重复)。5.3 现象Apriori结果全是单字词“盐”“油”“水”无实际价值原因min_support设太高或事务矩阵构建时没过滤停用词。解决在find_ingredient_pairs()前加停用词过滤STOPWORDS {盐, 油, 水, 料酒, 糖, 鸡精, 味精, 胡椒粉} basket basket.drop(columns[col for col in basket.columns if col in STOPWORDS], errorsignore)5.4 现象Flask本地运行正常部署到Linux服务器后图表空白原因Plotly CDN在部分企业内网被拦截或服务器时间不同步导致HTTPS证书失效。解决下载Plotly离线版放入static/js/目录改引用!-- 替换CDN链接 -- script src{{ url_for(static, filenamejs/plotly-2.24.1.min.js) }}/script同时检查服务器时间sudo ntpdate -s time.nist.gov。5.5 现象聚类结果每次运行都不一样无法复现原因KMeans随机初始化random_state没全局统一。解决在clustering.py顶部加全局seed并确保所有随机操作用同一seedimport numpy as np import random SEED 42 np.random.seed(SEED) random.seed(SEED) # 所有sklearn模型都传 random_stateSEED6. 进阶技巧用Docker一键部署看板 添加搜索框实时过滤6.1 Docker化部署三步打包彻底解决环境依赖本地跑通后用Docker封装成镜像避免“在我机器上好好的”问题。Dockerfile极简FROM python:3.9-slim WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt COPY . . EXPOSE 5000 CMD [gunicorn, --bind, 0.0.0.0:5000, --workers, 2, app:app]requirements.txt关键依赖Flask2.3.3 requests-html0.10.0 pandas2.1.3 plotly6.16.0 gunicorn21.2.0 alembic1.13.1构建命令docker build -t mt-food-dashboard . docker run -p 5000:5000 -v $(pwd)/data:/app/data mt-food-dashboard注意-v $(pwd)/data:/app/data将宿主机data/目录挂载到容器内确保SQLite数据库文件持久化。容器内路径必须和config.py中数据库路径一致sqlite:///data/app.db。6.2 前端增强添加搜索框实时过滤食材热力图templates/dashboard.html增加搜索框和JS逻辑实现输入即查input typetext idsearch-input placeholder搜索食材... oninputfilterHeatmap(this.value) div idheatmap stylewidth: 900px; height: 500px;/div script let allData []; // 全局缓存原始热力图数据 function loadHeatmap() { fetch(/api/associations/heatmap) .then(r r.json()) .then(data { allData data; renderHeatmap(data); }); } function filterHeatmap(keyword) { if (!keyword.trim()) return renderHeatmap(allData); const filtered allData.filter(item item.antecedents.includes(keyword) || item.consequents.includes(keyword) ); renderHeatmap(filtered); } function renderHeatmap(data) { // Plotly热力图代码略同4.2节结构 } /script后端新增接口/api/associations/heatmap返回完整规则列表非TOP20供前端自由筛选。6.3 数据验证技巧用pytest写三个必跑测试为防重构破坏核心逻辑我在tests/目录写了三个轻量测试# tests/test_scraper.py def test_ingredient_standardization(): assert standardize_ingredient(五花腩) 五花肉 assert standardize_ingredient(马铃薯) 土豆 # tests/test_association.py def test_apriori_min_support(): # 用小样本数据测试确保支持度阈值生效 sample_basket pd.DataFrame({ 鸡蛋: [1,1,0,0], 葱: [1,1,1,0], 盐: [1,1,1,1] }) freq apriori(sample_basket, min_support0.5, use_colnamesTrue) assert len(freq) 2 # 只有鸡蛋葱、葱盐满足0.5支持度 # tests/test_api.py def test_api_top10_returns_json(): app.config[TESTING] True client app.test_client() rv client.get(/api/ingredients/top10) assert rv.status_code 200 assert isinstance(rv.get_json(), list)运行命令pytest tests/ -v。从那以后我每次改cleaner.py或association.py都强制跑这三测——后悔药不如预防针管用。希望帮到你。本文还有配套的精品资源点击获取
阅读完成 · 觉得有帮助?