一、整体流程
配置参数 → 循环翻页 → 请求列表接口 → 解析 cards → 提取 mblog → 判断是否长文本 → 请求长文本接口 → 清洗文本 → 解析日期 → 去重 → 存入列表 → 写入 txt
注意:
整个爬虫只有一个主循环for page in range(1, MAX_PAGES + 1),每页调用一次fetch_page(page),然后遍历返回的cards。
二、请求网页列表接口
核心代码:
url = "https://m.weibo.cn/api/container/getIndex" params = { "type": "uid", "value": UID, "containerid": CONTAINERID, "page": page, } resp = requests.get(url, headers=HEADERS, params=params, timeout=20)解析:
UID是用户ID:2656274875。
CONTAINERID固定为107603 + UID,这是移动端用来标识“某用户列表”的容器ID。
page从1开始递增,每页大约返回10条数据。
必须带Cookie和移动端 User-Agent,否则会被重定向或返回空数据。
返回结构:
{ "ok": 1, "data": { "cards": [ { "card_type": 9, "mblog": { ... } }, ... ] } }注意:
只有card_type == 9的卡片才是原创/转发,其他类型(如广告、推荐)直接跳过。
三、提取正文与长文本判断
核心代码:
text = clean_html(mblog.get("text", "")) if mblog.get("isLongText", False): mblog_id = mblog.get("mblogid") or mblog.get("idstr") or mblog.get("id") long_text = fetch_long_text(mblog_id) if long_text: text = clean_html(long_text)解析:
列表页返回的text字段是摘要,如果数据超过140字,isLongText会为True,此时text被截断。
长文本接口需要的是网页的字符串ID,移动端字段是mblogid,但部分帖子没有这个字段,所以用 `idstr或id回退。
长文本接口地址是移动端专用:
url = "https://m.weibo.cn/statuses/extend" params = {"id": mblog_id}返回的完整内容在data.longTextContent里。
注意:
长文本请求要复用移动端HEADERS,不要改成weibo.com。
每篇长文之间必须time.sleep(LONG_TEXT_SLEEP),否则容易被限制。
四、日期解析
核心代码:
def parse_created_at(created_at): ...该网页返回的created_at格式不统一,常见有:
2025年8月23日
2025-08-23
08-23
今天 12:30
昨天 08:15
3小时前
刚刚
注意:
函数用正则逐个匹配,统一转成YYYY年M月D日。对于“今天/昨天/几分钟前”,直接取当前日期或前一天;无法解析的保留原样。
五、去重与翻页控制
去重:
seen_ids = set() ... mid = mblog.get("id") if not mid or mid in seen_ids: continue seen_ids.add(mid) ```每抓一条就把id加入集合,防止翻页时重复抓取。
翻页与频率:
time.sleep(SLEEP_SECONDS) # 每页间隔 5 秒 time.sleep(LONG_TEXT_SLEEP) # 每篇长文间隔 3 秒注意:
对未登录/频繁请求有严格限制,间隔太小会返回418或 403。fetch_page和fetch_long_text内部都有retry重试,失败后等待RETRY_WAIT秒再试。
六、文本清洗
核心代码:
def clean_html(text): text = re.sub(r"<br\s*/?>", "\n", text) text = re.sub(r"<.*?>", "", text) text = unescape(text) text = re.sub(r"[ \t]+", " ", text) text = re.sub(r"\n\s*\n+", "\n", text) return text.strip()注意:
网页正文里可能包含<a>、<br>、 等 HTML 标签和实体,清洗后只保留纯文本,方便写入 txt。
七、写入文件格式
核心代码:
with open(OUTPUT_FILE, "w", encoding="utf-8") as f: for text, date_str in all_posts: f.write(text + "\n") f.write(date_str + "\n") f.write("\n")最终文件的格式:
正文内容……
2026年9月23日
下一条正文……
2026年9月22日
注意:
每条之间空一行,方便后续按“内容+日期”成对读取。
八、关键配置参数
| 参数 | 作用 | 建议 |
| MAX_PAGES | 最多翻多少页 | 10 页约 100 条 |
| MAX_POSTS | 最多抓多少条 | 防止无限抓取 |
| SLEEP_SECONDS | 每页间隔 | 5 秒以上 |
| LONG_TEXT_SLEEP | 长文间隔 | 3 秒以上 |
| RETRY_WAIT | 失败等待 | 15 秒 |
| COOKIE | 登录凭证 | 必须包含SUB=和MLOGIN=1 |
九、常见问题定位
| 现象 | 原因 | 解决 |
| ok != 1 | Cookie 过期 | 重新获取 Cookie |
| 418/403 | 请求太快 | 增大SLEEP_SECONDS |
| mblogid=None | 字段缺失 | 已用idstr/id回退 |
| 长文本失败 | 接口或 ID 错误 | 确认用m.weibo.cn/statuses/extend |
| 只抓十几条 | 未登录或限制 | 确认MLOGIN=1且带SUB |
十、完整代码示例
# -*- coding: utf-8 -*- """ 仅供个人学习研究使用,请勿用于商业用途。 """ import requests import re import time from datetime import datetime, timedelta from html import unescape # ==================== 配置区 ==================== UID = "2656274875" # 央视新闻 uid CONTAINERID = f"107603{UID}" # 用户微博容器 id OUTPUT_FILE = "央视新闻微博数据.txt" MAX_PAGES = 10 # 最多抓取页数,每页约 10 条 MAX_POSTS = 100 # 最多抓取条数 SLEEP_SECONDS = 5 # 每页间隔秒数 LONG_TEXT_SLEEP = 3 # 每篇长文之间的间隔秒数 RETRY_WAIT = 15 # 请求失败后的等待秒数 # ★★★ 把你的完整 Cookie 粘贴到下面引号中间 ★★★ COOKIE = "WEIBOCN_FROM=1110006030; _T_WM=40926697456; SCF=AsoqZryMxPRtY0vje-p6M9UoyxUR2SLP87X9qe8XWmUhYBoJ3eo5JHUaU8w2vzlvfyGY1C6QGv48BXxl1P7j2Ug.; SUB=_2A25Ht_yZDeRhGe5N6VYR8izFzzqIHXVkzXBRrDV6PUJbktANLVrEkW1NdM5p0UXmpTVJMSIGTjJD9LhpSH5WV3o1; SUBP=0033WrSXqPxfM725Ws9jqgMF55529P9D9WWXaKG5dvjZrasvGZP.aKFi5NHD95QRe0zXehzE1KBcWs4DqcjMi--NiK.Xi-2Ri--ciKnRi-zN1heESh5Eeo.XSntt; SSOLoginState=1790151881; ALF=1792743881; MLOGIN=1; XSRF-TOKEN=5cb1b8" # ==================== 请求头 ==================== HEADERS = { "User-Agent": ( "Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) " "AppleWebKit/605.1.15 (KHTML, like Gecko) " "Version/16.0 Mobile/15E148 Safari/604.1" ), "Referer": f"https://m.weibo.cn/u/{UID}", "X-Requested-With": "XMLHttpRequest", "Accept": "application/json, text/plain, */*", "MWeibo-Pwa": "1", "Cookie": COOKIE, } # ==================== 工具函数 ==================== def clean_html(text): """去除 HTML 标签,保留纯文本""" if not text: return "" text = re.sub(r"<br\s*/?>", "\n", text) text = re.sub(r"<.*?>", "", text) text = unescape(text) text = re.sub(r"[ \t]+", " ", text) text = re.sub(r"\n\s*\n+", "\n", text) return text.strip() def parse_created_at(created_at): """把微博的 created_at 转成 YYYY年M月D日 格式""" if not created_at: return "" s = created_at.strip() now = datetime.now() m = re.match(r"(\d{4})年(\d{1,2})月(\d{1,2})日", s) if m: return f"{m.group(1)}年{int(m.group(2))}月{int(m.group(3))}日" m = re.match(r"(\d{4})-(\d{1,2})-(\d{1,2})", s) if m: return f"{m.group(1)}年{int(m.group(2))}月{int(m.group(3))}日" m = re.match(r"^(\d{1,2})-(\d{1,2})$", s) if m: month, day = int(m.group(1)), int(m.group(2)) year = now.year if month > now.month: year -= 1 return f"{year}年{month}月{day}日" if s.startswith("今天"): return f"{now.year}年{now.month}月{now.day}日" if s.startswith("昨天"): y = now - timedelta(days=1) return f"{y.year}年{y.month}月{y.day}日" if "刚刚" in s or "分钟前" in s or "小时前" in s: return f"{now.year}年{now.month}月{now.day}日" return s def fetch_page(page, retry=3): """请求一页微博列表,失败自动重试""" url = "https://m.weibo.cn/api/container/getIndex" params = { "type": "uid", "value": UID, "containerid": CONTAINERID, "page": page, } for attempt in range(1, retry + 1): try: resp = requests.get(url, headers=HEADERS, params=params, timeout=20) if resp.status_code == 200: data = resp.json() if data.get("ok") == 1: return data else: print(f" 第 {page} 页返回 ok != 1,msg={data.get('msg', '')}") return None elif resp.status_code in (418, 403): print(f" 第 {page} 页被限制,状态码 {resp.status_code},等待 {RETRY_WAIT} 秒后重试({attempt}/{retry})...") time.sleep(RETRY_WAIT) else: print(f" 第 {page} 页请求失败,状态码 {resp.status_code},等待 {RETRY_WAIT} 秒后重试({attempt}/{retry})...") time.sleep(RETRY_WAIT) except Exception as e: print(f" 第 {page} 页请求异常:{e},等待 {RETRY_WAIT} 秒后重试({attempt}/{retry})...") time.sleep(RETRY_WAIT) print(f" 第 {page} 页重试 {retry} 次仍失败,跳过。") return None def fetch_long_text(mblog_id, retry=3): """请求移动端长文本接口,返回完整内容""" if not mblog_id: return "" # ★ 注意:这里改为移动端专用的长文本接口 url = "https://m.weibo.cn/statuses/extend" params = {"id": mblog_id} for attempt in range(1, retry + 1): try: # 使用移动端统一的 HEADERS resp = requests.get(url, headers=HEADERS, params=params, timeout=20) if resp.status_code == 200: data = resp.json() if data.get("ok") == 1: return data.get("data", {}).get("longTextContent", "") else: print(f" 长文本接口返回 ok != 1,msg={data.get('msg', '')}") return "" elif resp.status_code in (418, 403): print(f" 长文本被限制,状态码 {resp.status_code},等待 {RETRY_WAIT} 秒后重试({attempt}/{retry})...") time.sleep(RETRY_WAIT) else: print(f" 长文本请求失败,状态码 {resp.status_code},等待 {RETRY_WAIT} 秒后重试({attempt}/{retry})...") time.sleep(RETRY_WAIT) except Exception as e: print(f" 长文本请求异常:{e},等待 {RETRY_WAIT} 秒后重试({attempt}/{retry})...") time.sleep(RETRY_WAIT) print(f" 长文本重试 {retry} 次仍失败,跳过。") return "" # ==================== 主流程 ==================== def main(): all_posts = [] seen_ids = set() print("=" * 60) print("开始抓取央视新闻微博数据(含长文本完整内容)...") print("=" * 60) for page in range(1, MAX_PAGES + 1): print(f"\n>>> 正在抓取第 {page} 页...") data = fetch_page(page) if not data: print(" 本页无数据,停止抓取。") break cards = data.get("data", {}).get("cards", []) if not cards: print(" 没有更多卡片,停止抓取。") break page_posts = [] for card in cards: if card.get("card_type") != 9: continue mblog = card.get("mblog") if not mblog: continue mid = mblog.get("id") if not mid or mid in seen_ids: continue seen_ids.add(mid) # 先取列表页的摘要文本 text = clean_html(mblog.get("text", "")) # 判断是否为长文本 if mblog.get("isLongText", False): # ★ 核心修复:字段回退逻辑 # 优先取 mblogid,若为空则取 idstr,最后取 id mblog_id = mblog.get("mblogid") or mblog.get("idstr") or mblog.get("id") print(f" 检测到长文本,正在获取完整内容 (id={mblog_id})...") long_text = fetch_long_text(mblog_id) if long_text: text = clean_html(long_text) print(f" 长文本获取成功,长度 {len(text)} 字。") else: print(f" 长文本获取失败,保留摘要。") time.sleep(LONG_TEXT_SLEEP) created_at_raw = mblog.get("created_at", "") date_str = parse_created_at(created_at_raw) if not text: continue page_posts.append((text, date_str)) if not page_posts: print(" 本页没有有效微博,停止抓取。") break all_posts.extend(page_posts) print(f" 本页获取 {len(page_posts)} 条,累计 {len(all_posts)} 条。") if len(all_posts) >= MAX_POSTS: print(f"\n已达到最大抓取条数 {MAX_POSTS},停止。") break time.sleep(SLEEP_SECONDS) # ==================== 写入文件 ==================== with open(OUTPUT_FILE, "w", encoding="utf-8") as f: for text, date_str in all_posts: f.write(text + "\n") f.write(date_str + "\n") f.write("\n") print("\n" + "=" * 60) print(f"抓取完成!共 {len(all_posts)} 条微博。") print(f"已保存到:{OUTPUT_FILE}") print("=" * 60) if __name__ == "__main__": main()