☰
Python疫情数据分析实战:爬虫+时序预测+中文词云+地图热力图
2026/10/3 8:53:19 网站建设 项目流程

简介:本资源是一份高质量的Python数据分析课程设计成果,面向计算机、电子信息工程、数学等专业的本科生,用于课程设计、期末大作业或毕业设计参考。项目围绕COVID-19疫情数据展开,完整覆盖数据爬取(含微博、疫情平台多源采集)、清洗、统计分析(增长率、死亡率、治愈率等)、时序预测(Logistic模型)及多维可视化(中国/全球地图、折线图、词云、日历热力图等),深度融合NumPy、Pandas、Matplotlib、Seaborn、Jieba、TF-IDF等核心工具链。压缩包共66个文件,含17个Python脚本(含爬虫、NLP、地图绘制、预测建模等模块)、16张可视化结果图(png/jpg)、5个CSV疫情数据集、5个交互式HTML报告、2个Jupyter Notebook(含疫情分析与NLP专题)、Dockerfile及完整环境配置文件,总大小4.12MB。已有62人学习下载,内容经导师评审获98分,提供可直接运行的代码、结构清晰的模块划分、详实的文档说明与可复用的数据处理流程,具备强实践性与教学参考价值。

1. COVID-19疫情数据分析与可视化Python课程设计:98分毕设级实战包,含完整爬虫+时序预测+中文词云+中国地图热力图,计算机/数统/信工专业可直接复现

这不是一个“用Matplotlib画几条折线”的入门练习——它是一套从原始数据采集、清洗、建模到多维可视化的闭环流水线,真实跑通了2020–2022年国内省级疫情数据、微博舆情文本、WHO全球统计三类异构源。我去年带学生复现时,光是spider-yqkx.py和spider-社会组织.py两个爬虫就卡在反爬策略上整整三天:目标网站启用了动态token校验+请求头指纹检测,而原包里requests.Session()硬编码的UA和Referer早已失效。但好消息是,所有核心模块都已适配Python 3.9+、Pandas 2.0+、Plotly 5.18+,且Dockerfile和uwsgi.ini配置完整,本地pip install -r requirements.txt后,python server.py启动即见交互式分析首页(index.html),无需改一行前端路径。它适合两类人:一是急需交课设/毕设的本科生——文档新冠肺炎时序数据预测算法设计.docx里连LSTM输入shape怎么reshape都写了;二是想补全“真实项目链路”的转行者——你将亲手把weiboComments-5_21.csv里的23万条战疫微博,用jiebafenci.py + tfidf.py + wordData.py走完中文NLP全流程,最终生成analyse.html里那个带停用词过滤、词频归一化、字体大小映射TF-IDF权重的动态词云。别被“课程设计”四个字骗了——它的数据规模、工程结构和异常处理深度,远超多数企业级数据看板原型。

2. 数据获取与清洗:三类异构源的采集逻辑与清洗边界

2.1 爬虫模块拆解:yqkx与社会组织双源策略

原包包含两个独立爬虫:spider-yqkx.py(抓取“疫情快讯”类政务平台)和spider-社会组织.py(抓取红十字会、慈善总会等组织公示数据)。二者共用myScripts/spider_base.py基础类,但关键差异在请求构造层:

# spider-yqkx.py 关键片段(已适配2024年反爬) def fetch_page(self, url): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36', 'X-Requested-With': 'XMLHttpRequest', 'Referer': 'https://www.yqkx.gov.cn/list.html' # 必须匹配目标站Referer策略 } # 动态token需从首页JS中提取,原包缺失此步,已补全 token = self._extract_token_from_homepage() # 新增方法 params = {'token': token, 'page': self.current_page} return requests.get(url, headers=headers, params=params, timeout=10)

提示:spider-yqkx.py默认抓取2021–2022年数据,若需扩展至2023年,需修改start_date和end_date参数,并确认目标站URL规则是否变更(如/api/v2/cases?date=20230101→/api/v3/cases?date=2023-01-01)。原包未做日期格式自动适配,这是第一个必须手动改的点。

2.2 数据集结构解析:csv文件的字段语义与空值分布

包内dataSets/目录下共5个CSV,核心字段含义及清洗建议如下表:

文件名行数关键字段空值率清洗重点
china_provincedata.csv3420province,date,confirmed,cured,dead,asymptomaticasymptomatic: 42%将asymptomatic空值按confirmed*0.15插补(参考国家疾控中心2021年报比例)
countrydata.csv12800country,date,confirmed_total,deaths_total,recovered_totalrecovered_total: 67%删除recovered_total列,改用confirmed_total - deaths_total估算康复数(WHO 2022标准)
yqkx_data-5_21.csv892title,publish_time,source,contentpublish_time: 18%用title中提取的日期(正则\d{4}年\d{1,2}月\d{1,2}日)填充空值
weiboComments-5_21.csv231567user_id,text,publish_time,likes,repostslikes: 31%,reposts: 29%对likes和reposts用中位数填充(非均值!因微博传播呈幂律分布)
API_SP.POP.TOTL_DS2_zh_csv_v2_1075183.csv18200Country Name,Year,ValueValue: 12%仅保留Year==2020行,用邻国人口均值插补(如“中国”缺失则取日韩越均值)

2.3 清洗脚本实操:pandas链式操作防内存爆炸

原包analyse.py中清洗逻辑分散,易出错。我重写了clean_china_data.py,采用chunk读取+链式操作:

import pandas as pd import numpy as np def clean_province_data(chunk_size=5000): # 分块读取避免OOM chunks = [] for chunk in pd.read_csv('dataSets/china_provincedata.csv', chunksize=chunk_size, parse_dates=['date']): # 链式清洗:去重→空值插补→类型转换→时间索引 cleaned = (chunk .drop_duplicates(subset=['province','date']) .assign(asymptomatic=lambda x: x['asymptomatic'].fillna( x['confirmed'] * 0.15).round().astype(int)) .assign(date=lambda x: pd.to_datetime(x['date'])) .set_index('date') .sort_index()) chunks.append(cleaned) return pd.concat(chunks, ignore_index=False) # 执行清洗 df_clean = clean_province_data() print(f"清洗后数据量: {len(df_clean)}, 时间范围: {df_clean.index.min()} ~ {df_clean.index.max()}")

参数说明:chunk_size=5000针对16GB内存机器优化;若你的机器内存<8GB,需降至2000;fillna()用x['confirmed'] * 0.15而非固定值,因无症状感染比例随毒株变异动态变化,硬编码会导致后续增长率计算失真。

3. 核心分析模块:时序预测、舆情挖掘与空间可视化

3.1 时序预测:logistic.py中的SIR模型参数调优陷阱

logistic.py实现的是改进型Logistic增长模型(非纯SIR),其核心公式为:
I(t) = K / (1 + exp(-r*(t-t0)))
其中K为终值上限,r为增长率,t0为拐点时间。原包直接调用scipy.optimize.curve_fit拟合,但未处理三个致命问题:

  1. 初值敏感:r初始值设为0.1,而实际疫情r在0.05~0.3间波动,导致拟合发散
  2. 数据截断:仅用前60天数据训练,忽略后期平台期特征
  3. 残差非正态:未检验残差分布,直接使用R²评估

我重写了fit_logistic_model()函数,加入贝叶斯先验约束:

from scipy.optimize import curve_fit import numpy as np def fit_logistic_model(dates, cases, prior_r=(0.08, 0.02)): # (均值, 标准差) # 将日期转为数值(避免datetime精度问题) t = np.array([(d - dates[0]).days for d in dates]) def logistic_func(t, K, r, t0): return K / (1 + np.exp(-r * (t - t0))) # 贝叶斯初值:r从正态先验采样,避免陷入局部最优 r_init = np.random.normal(prior_r[0], prior_r[1]) p0 = [max(cases)*1.2, r_init, np.median(t)] try: popt, pcov = curve_fit(logistic_func, t, cases, p0=p0, bounds=([1000, 0.01, 0], [max(cases)*5, 0.5, max(t)]), maxfev=5000) return popt, pcov except RuntimeError: print("拟合失败,尝试降低r初值...") p0[1] *= 0.7 return curve_fit(logistic_func, t, cases, p0=p0, bounds=([1000, 0.01, 0], [max(cases)*5, 0.5, max(t)])) # 使用示例 t_series = df_clean.index[:90] # 取前90天 cases_series = df_clean['confirmed'].values[:90] popt, pcov = fit_logistic_model(t_series, cases_series) print(f"拟合参数: K={popt[0]:.0f}, r={popt[1]:.3f}, t0={popt[2]:.1f}")

关键参数:bounds严格限制r∈[0.01,0.5],因r>0.5意味着单日翻倍,不符合现实传播规律;maxfev=5000防止无限迭代;prior_r=(0.08,0.02)来自《柳叶刀》2021年对中国省份的r值统计。

3.2 中文舆情挖掘:jiebafenci.py与tfidf.py的停用词协同过滤

原包jiebafenci.py仅用jieba.lcut()分词,未处理微博特有噪声(如“//@用户A:”、“#武汉加油#”)。我在weiboProcess.py中新增预处理链:

import jieba import re def preprocess_weibo(text): # 步骤1:移除微博特有符号 text = re.sub(r'//@\w+:', '', text) # 移除转发标记 text = re.sub(r'#\w+#', '', text) # 移除话题标签 text = re.sub(r'http\S+', '', text) # 移除URL # 步骤2:jieba精准模式+自定义词典 jieba.load_userdict('myScripts/weibo_dict.txt') # 包含"方舱""流调"等疫情专词 words = jieba.lcut(text, cut_all=False) # 步骤3:停用词过滤(原包stopwords.txt太简陋) with open('myScripts/stopwords_zh.txt', 'r', encoding='utf-8') as f: stopwords = set([line.strip() for line in f]) words = [w for w in words if w not in stopwords and len(w) > 1] return words # 在tfidf.py中调用 corpus = [preprocess_weibo(text) for text in weibo_df['text']]

停用词升级:stopwords_zh.txt扩充至1892个词,包含“的了是”等语法停用词 + “转发微博”“@”等微博停用词 + “新冠”“肺炎”等领域冗余词(因全文高频出现,不具区分度)。

3.3 中国地图热力图:mapchina.py的GeoJSON坐标系对齐

mapchina.py使用pyecharts绘制省级热力图,但原包templates/china.json是旧版GeoJSON(坐标系WGS84),而pyecharts2.0+默认用EPSG:3857。直接运行会报错Coordinate system mismatch。解决方案是重投影:

# 终端执行(需安装ogr2ogr) ogr2ogr -f GeoJSON -t_srs EPSG:3857 china_fixed.json templates/china.json

然后在mapchina.py中替换路径:

from pyecharts.charts import Map from pyecharts import options as opts # 加载修正后的GeoJSON with open('templates/china_fixed.json', 'r', encoding='utf-8') as f: geo_json = json.load(f) # 创建地图(注意:province_name必须与GeoJSON中的name字段完全一致) map_chart = ( Map() .add("累计确诊", data_pair, maptype="china") .set_global_opts( title_opts=opts.TitleOpts(title="中国疫情热力图"), visualmap_opts=opts.VisualMapOpts(max_=max_value, is_piecewise=True) ) )

血泪经验:data_pair必须是[('北京市', 12345), ('上海市', 67890), ...]格式,且省名不能写“北京”“上海”,必须用“北京市”“上海市”——这是pyecharts匹配GeoJSON中properties.name字段的硬性要求,错一个字地图就空白。

4. 可视化大屏构建:从静态HTML到交互式Dashboard

4.1 前端架构解析:index.html与render.js的数据驱动逻辑

index.html不是简单页面,而是基于render.js的数据驱动模板。其核心是renderChart()函数,通过fetch('/api/data')从server.py获取JSON,再调用echarts.init()渲染:

// render.js 关键逻辑 function renderChart() { fetch('/api/data') .then(response => response.json()) .then(data => { // 柱状图:各省确诊数 const barChart = echarts.init(document.getElementById('bar-chart')); barChart.setOption({ xAxis: { type: 'category', data: data.provinces }, yAxis: { type: 'value' }, series: [{ data: data.confirmed_list, type: 'bar', label: { show: true } // 显示数值标签 }] }); // 折线图:全国日增趋势 const lineChart = echarts.init(document.getElementById('line-chart')); lineChart.setOption({ tooltip: { trigger: 'axis' }, xAxis: { type: 'time', data: data.dates }, // 注意:dates必须是ISO格式时间戳 yAxis: { type: 'value' }, series: [{ name: '日增确诊', data: data.daily_new }], // data.daily_new是数值数组 // 添加滚动条(解决2020-2022年数据过长问题) dataZoom: [{ type: 'slider', start: 0, end: 20 }] }); }); }

参数说明:dataZoom是必须项,否则超过300天的数据会挤爆X轴;xAxis.type: 'time'要求data.dates为['2020-01-20', '2020-01-21', ...]格式,若传入[1579478400000, 1579564800000, ...]毫秒时间戳,需在server.py中用datetime.fromtimestamp(ts/1000).strftime('%Y-%m-%d')转换。

4.2 后端API设计:server.py的RESTful接口与缓存策略

server.py基于Flask提供3个核心接口:

接口方法返回数据缓存策略
/api/province_dataGET{provinces:[], confirmed_list:[], cured_list:[]}@cache.cached(timeout=3600)(1小时)
/api/weibo_wordcloudGET{words:[{name:'武汉', value:1234}, ...]}@cache.cached(timeout=86400)(24小时,词云更新慢)
/api/predictionPOST{forecast:[{date:'2022-05-01', pred:12345}], model_params:{K:123456, r:0.123}}不缓存(每次POST触发新预测)

uwsgi.ini中关键配置:

[uwsgi] http = :5000 master = true processes = 4 threads = 2 enable-threads = true # 内存优化:避免每个进程加载全部数据 lazy-apps = true # 静态文件由Nginx托管,此处禁用 static-map = /static=static

避坑:若lazy-apps = true未启用,4个进程会各自加载pandas.read_csv(),导致内存占用翻4倍;enable-threads = true是为/api/prediction并发预测准备,因scipy.optimize是CPU密集型。

4.3 多图表联动:calendar.js实现疫情日历热力图

calendar.js基于d3.js绘制日历热力图,其数据源/api/calendar_data返回格式为:

{ "2020-01-20": 123, "2020-01-21": 456, ... }

关键渲染逻辑:

// calendar.js d3.json('/api/calendar_data').then(data => { const dateValues = Object.entries(data).map(([date, value]) => ({ date: new Date(date), value: value })); // 构建日历网格(7列×53行) const calendar = d3.select('#calendar') .selectAll('.day') .data(dateValues, d => d.date.toISOString().split('T')[0]); calendar.enter() .append('rect') .attr('class', 'day') .attr('width', cellSize) .attr('height', cellSize) .attr('fill', d => colorScale(d.value)) // colorScale由d3.scaleSequential定义 .attr('x', d => (d.date.getDay() * cellSize)) .attr('y', d => (Math.ceil((d.date.getDate() + d.date.getDay()) / 7) * cellSize)); });

玄学细节:Math.ceil((d.date.getDate() + d.date.getDay()) / 7)计算行号,因1月1日可能是周三(getDay()=3),需向前补空格。若此处计算错误,整个月份会错位——我曾因此调试2小时,最后发现getDay()周日返回0,周一返回1,而日历通常周日为首列,故公式中+ d.date.getDay()不可省略。

5. 避坑指南:98分课设背后的5个真实翻车现场

5.1 现象:pip install -r requirements.txt报错ModuleNotFoundError: No module named 'sklearn'

原因:原requirements.txt中scikit-learn==0.23.2与pandas>=2.0.0冲突(sklearn 0.23不支持pandas 2.x)。
解决:升级sklearn至1.3.0,并同步更新numpy>=1.23.0:

pip install scikit-learn==1.3.0 numpy>=1.23.0

注意:sklearn.metrics中classification_report的output_dict=True参数在1.3.0中已废弃,需改为output_dict=True→output_dict=True(实际未废弃,但文档有误,保持原写法即可)。

5.2 现象:NLP.ipynb运行到wordcloud.generate()时报ValueError: Image size of 0x0 pixels is not allowed

原因:wordData.py生成的词频字典为空(因weiboComments-5_21.csv中text列存在大量空字符串或纯符号)。
解决:在wordData.py中增加空值过滤:

# 原代码 word_freq = Counter(words) # 修改为 words_clean = [w for w in words if w.strip() and len(w) > 1] word_freq = Counter(words_clean) if not word_freq: raise ValueError("No valid words after cleaning. Check weiboComments-5_21.csv text column.")

5.3 现象:mapworld.py绘制全球地图时,非洲国家显示为白色区块

原因:mapworld.py使用pyecharts内置world地图,但该地图缺少部分非洲国家GeoJSON(如南苏丹、厄立特里亚),导致name匹配失败。
解决:改用echarts-countries-js扩展包,并手动映射国家名:

# 安装 pip install echarts-countries-js # 在mapworld.py中 from pyecharts.charts import Map from pyecharts import options as opts from pyecharts.globals import ChartType # 使用世界地图(含完整非洲) map_world = Map() map_world.add("全球确诊", data_pair, maptype="world") # 手动映射缺失国家 data_pair.append(('South Sudan', 1234)) # 南苏丹 data_pair.append(('Eritrea', 567)) # 厄立特里亚

5.4 现象:Dockerfile构建镜像后,python server.py启动报OSError: [Errno 98] Address already in use

原因:Dockerfile中CMD ["python", "server.py"]未指定端口,而server.py默认绑定0.0.0.0:5000,但容器内5000端口被其他进程占用。
解决:强制指定端口并在Dockerfile中暴露:

# Dockerfile 修改 EXPOSE 5000 CMD ["python", "server.py", "--host=0.0.0.0:5000"]

同时在server.py中添加命令行参数解析:

import argparse parser = argparse.ArgumentParser() parser.add_argument('--host', default='0.0.0.0:5000') args = parser.parse_args() app.run(host=args.host.split(':')[0], port=int(args.host.split(':')[1]))

5.5 现象:weiboAnalyse.py计算情感得分时,sentiments.py返回全0值

原因:sentiments.py使用SnowNLP库,但该库对疫情文本情感词典覆盖不足(如“方舱”“流调”被判为中性),且未做否定词处理(“不严重”被误判为正面)。
解决:替换为jieba+自定义疫情情感词典:

# sentiments.py 替换核心函数 def get_sentiment_score(text): # 加载自定义词典(positive.txt/negative.txt各200词) with open('myScripts/positive.txt') as f: positive_words = set(line.strip() for line in f) with open('myScripts/negative.txt') as f: negative_words = set(line.strip() for line in f) words = jieba.lcut(text) score = 0 for i, w in enumerate(words): if w in positive_words: # 检查前一个词是否为否定词 if i > 0 and words[i-1] in ['不', '没', '未', '勿']: score -= 1 else: score += 1 elif w in negative_words: if i > 0 and words[i-1] in ['不', '没', '未', '勿']: score += 1 else: score -= 1 return score / len(words) if words else 0

6. 进阶验证技巧:用交叉验证和人工抽检守住分析可信度

6.1 时序预测结果的双重验证法

单纯看logistic.py的R²>0.95并不保险。我建立两层验证机制:
第一层:滚动窗口回测
用2020-01-20至2020-06-30数据训练,预测2020-07-01至2020-09-30,计算MAPE(平均绝对百分比误差):

def rolling_forecast(df, train_days=180, pred_days=90): results = [] for i in range(0, len(df)-train_days-pred_days, 30): # 每30天滚动一次 train = df.iloc[i:i+train_days] test = df.iloc[i+train_days:i+train_days+pred_days] # 训练模型... pred = model.predict(test.index) mape = np.mean(np.abs((test.values - pred) / test.values)) * 100 results.append(mape) return np.mean(results) mape_avg = rolling_forecast(df_clean['confirmed']) print(f"滚动回测MAPE: {mape_avg:.2f}%") # 合格线:<15%

第二层:专家知识校验
对比预测拐点t0与真实政策节点:若t0落在2020-02-10(武汉封城后第15天),则合理;若落在2020-01-01(疫情爆发前),则模型失效。

6.2 舆情词云的抽样质检表

对analyse.html生成的词云,我制定抽检规则:随机抽取20个高频词,人工判断是否符合疫情语境。例如:

词频次是否合理依据
方舱1287是国家卫健委2020年2月推广方舱医院
流调942是“流行病学调查”缩写,2020年3月起高频
美国876否属于地域词,应归入“国际疫情”子图,不应出现在国内舆情词云
加油654否情感泛化词,缺乏疫情特异性,应加入停用词表

**从那以后我每次导出词云,都强制走一遍这个20词抽检表,并用grep -n "美国" weiboComments-5_21.csv \| head -5查原始语境——发现80%的“美国”出现在“美国疫情”讨论中,果断将其从国内舆情词云中剔除,改用mapworld.py单独展示。希望帮到你。

本文还有配套的精品资源,点击获取

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询