1. 竞彩网 NBA 受注数据抓取,难在哪
竞彩网 NBA 受注数据是很多做体育数据看板、赔率监控、赛前分析的同学绕不开的一类数据源。它不像普通新闻页那样一页 HTML 就能扒完,而是分成两层:第一层是比赛列表接口,返回当天或指定联赛的全部场次;第二层是每场比赛的详情接口,返回主客队场均得失分、近期胜负、交锋战绩、胜率这些受注参考指标。两层数据靠matchId串起来,必须先在列表里拿到 id,再带着 id 去请求详情,最后把两份数据合并成一条完整记录。
这个场景适合已经会写基础 Scrapy 爬虫、想练手「多级 Request + items 字段拼装 + json 输出」的同学。核心检索词就是 scrapy、爬虫、json、Request、items,本文会围绕这几个点,把竞彩网 NBA 受注数据的完整抓取流程走一遍,包括 settings 配置、items 定义、两层请求的 callback 写法、json 落盘,以及用 TaoToken 统一 Key 做接口调试的配置骨架。
我试过直接拿浏览器里复制的 URL 去跑,结果第二层详情请求全部 403,后来才发现是请求头里少了 Referer 和 UA,加上之后才稳定。下面按「先跑通列表、再拼详情、最后输出 json」的顺序来,每一步都给可复制的代码。
2. TaoToken 前置:统一 Key 与接口调试骨架
在写爬虫的过程中,经常需要临时调一个模型接口来解析字段含义、生成 items 字段映射,或者把抓到的 json 丢给模型做结构化校验。如果每个脚本都单独配一套 Key,管理起来很乱。TaoToken 提供的是统一 Key 的方式,一个 Key 走多个模型,适合放在爬虫项目的配置里做辅助调试。
官网地址是 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,API 入口是 https://taotoken.net/api ,注意 API 地址不带 UTM 参数。你可以在控制台里创建 Key,然后把它写进项目的settings.py或者单独的config.py,爬虫主流程和调试脚本共用同一个 Key。
具体操作路径:先打开控制台 https://taotoken.net/console?utm_source=taotoken_aicg_blog_end&utm_content=console&utm_campaign=rewrite ,在 API Keys 页面 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api-keys&utm_campaign=rewrite 新建一个 Key,复制出来。如果你只是想验证某个模型能不能正确解析竞彩字段,可以直接用模型对话页 https://taotoken.net/models?utm_source=taotoken_aicg_blog_end&utm_content=models&utm_campaign=rewrite 试一下,把一段 json 贴进去问「这几个字段分别代表什么」。长期做编码和 Agent 任务的话,Coding Plan 页面 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding-plan&utm_campaign=rewrite 有对应的套餐说明。
注意:TaoToken 在这里的角色是「接口调试与字段解析的辅助工具」,不是爬虫的代理层,也不要用它去绕过目标站点的访问限制。爬虫本身的请求头、频率控制还是要自己做好。
配置骨架可以这样写,放在项目根目录的config.py:
# config.py TAOTOKEN_API_BASE = "https://taotoken.net/api" TAOTOKEN_API_KEY = "你的Key" # 数据库配置 DB_CONN = { "host": "127.0.0.1", "port": 3306, "user": "dev", "password": "123456", "database": "crawler", "charset": "utf8mb4", }然后在settings.py里引入,爬虫主流程和调试脚本都从这里读,避免 Key 散落在多个文件里。
3. 可复制配置:items 定义与 settings.json 输出
先建 Scrapy 项目,假设项目名叫lottery_crawls:
scrapy startproject lottery_crawls cd lottery_crawls scrapy genspider bask_station webapi.sporttery.cn3.1 items.py 字段定义
竞彩 NBA 受注数据字段比较多,列表接口给一部分,详情接口给另一部分,items 要能装下两层合并后的结果:
# lottery_crawls/items.py import scrapy class BaskStation(scrapy.Item): no_str = scrapy.Field() # 场次编号 league = scrapy.Field() # 联赛简称 datetime = scrapy.Field() # 比赛时间 home = scrapy.Field() # 主队简称 away = scrapy.Field() # 客队简称 a_sf = scrapy.Field() # 客队近期胜负 h_sf = scrapy.Field() # 主队近期胜负 battle = scrapy.Field() # 交锋战绩 a_rate = scrapy.Field() # 客队胜率 h_rate = scrapy.Field() # 主队胜率 a_ds = scrapy.Field() # 客队场均得失分 h_ds = scrapy.Field() # 主队场均得失分3.2 settings.py 关键配置
settings.py里要改的地方不多,但每一项都影响能不能跑通。重点是ROBOTSTXT_OBEY关掉、DEFAULT_REQUEST_HEADERS补上、ITEM_PIPELINES打开、FEEDS配 json 输出:
# lottery_crawls/settings.py BOT_NAME = "lottery_crawls" SPIDER_MODULES = ["lottery_crawls.spiders"] NEWSPIDER_MODULE = "lottery_crawls.spiders" ROBOTSTXT_OBEY = False DOWNLOAD_DELAY = 1 CONCURRENT_REQUESTS = 4 DEFAULT_REQUEST_HEADERS = { "Accept": "application/json, text/plain, */*", "Accept-Language": "zh-CN,zh;q=0.9", "Referer": "https://www.sporttery.cn/", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) " "AppleWebKit/537.36 (KHTML, like Gecko) " "Chrome/120.0.0.0 Safari/537.36", } ITEM_PIPELINES = { "lottery_crawls.pipelines.BaskStationPipeline": 300, } # json 输出,按时间戳命名,避免覆盖 from datetime import datetime FEEDS = { f"output/bask_{datetime.now().strftime('%Y%m%d_%H%M%S')}.json": { "format": "json", "encoding": "utf8", "indent": 2, "ensure_ascii": False, }, } REQUEST_FINGERPRINTER_IMPLEMENTATION = "2.7" TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"FEEDS里用ensure_ascii: False是为了让中文正常显示,不然 json 里全是\uXXXX,校验的时候很难看。
3.3 两层 Request 的 spider 写法
列表接口返回的是按日期分组的matchInfoList,每组下面有subMatchList,每个子项就是一场比赛。拿到matchId后再发详情请求,用meta把列表阶段拼好的 item 带过去:
# lottery_crawls/spiders/bask_station.py import json import scrapy from scrapy import Request from scrapy.http import TextResponse from lottery_crawls.items import BaskStation class BaskStationSpider(scrapy.Spider): name = "bask_station" allowed_domains = ["webapi.sporttery.cn"] def start_requests(self): url = ("https://webapi.sporttery.cn/gateway/jc/basketball/" "getMatchListV1.qry?clientCode=3001&leagueId=1") yield Request(url=url, callback=self.parse, dont_filter=True) def parse(self, response: TextResponse, **kwargs): data = json.loads(response.text) date_group = data["value"]["matchInfoList"] for group in date_group: for item in group["subMatchList"]: basketball_item = BaskStation() basketball_item["no_str"] = item["matchNumStr"] basketball_item["datetime"] = ( str(item["matchDate"])[5:16] + " " + item["matchTime"] ) basketball_item["home"] = item["homeTeamAbbName"] basketball_item["away"] = item["awayTeamAbbName"] basketball_item["league"] = item["leagueAbbName"] match_id = item["matchId"] detail_url = ( "https://webapi.sporttery.cn/gateway/uniform/basketball/" f"getMatchFeatureV1.qry?termLimits=10&gmMatchId={match_id}" ) yield Request( url=detail_url, meta={"key": basketball_item}, callback=self.data_analysis, dont_filter=True, ) def data_analysis(self, response: TextResponse): item = response.meta["key"] value = json.loads(response.text)["value"] a_d = value["scoreAvg"]["awayGoalAvgCnt"] a_s = value["lossScoreAvg"]["awayLossGoalAvgCnt"] item["a_ds"] = f"{a_d}/{a_s}" h_d = value["scoreAvg"]["homeGoalAvgCnt"] h_s = value["lossScoreAvg"]["homeLossGoalAvgCnt"] item["h_ds"] = f"{h_d}/{h_s}" a_win = str(value["eachHomeAway"]["awayWinGoalMatchCnt"]) a_lose = str(value["eachHomeAway"]["awayLossGoalMatchCnt"]) item["a_sf"] = f"{a_win}胜/{a_lose}负" h_win = str(value["eachHomeAway"]["homeWinGoalMatchCnt"]) h_lose = str(value["eachHomeAway"]["homeLossGoalMatchCnt"]) item["h_sf"] = f"{h_win}胜/{h_lose}负" item["a_rate"] = value["eachHomeAway"]["awayScoreRatio"] + "%" item["h_rate"] = value["eachHomeAway"]["homeScoreRatio"] + "%" home_name = value["homeTeamShortName"] b_win = str(value["last"]["homeWinGoalMatchCnt"]) b_lose = str(value["last"]["homeLossGoalMatchCnt"]) item["battle"] = f"主队{home_name}{b_win}胜{b_lose}负" yield item这里有个容易踩的点:callback只能写函数名字符串或者函数对象,不能写成callback=self.data_analysis()带括号,带括号就变成调用结果了,Scrapy 会报TypeError。
3.4 pipeline 落库与 json 双写
如果只是要 json 输出,FEEDS已经够了,pipeline 可以只做数据清洗。但实际项目里通常还要落库,这里给一个 pymysql 的管道,和 json 输出并存:
# lottery_crawls/pipelines.py import logging import pymysql from config import DB_CONN class BaskStationPipeline: def __init__(self): self.table_name = "game_baskstation" self.conn = pymysql.connect( host=DB_CONN["host"], port=DB_CONN["port"], user=DB_CONN["user"], password=DB_CONN["password"], database=DB_CONN["database"], charset=DB_CONN["charset"], ) self.cursor = self.conn.cursor() def close_spider(self, spider): self.cursor.close() self.conn.close() def process_item(self, item, spider): sql = ( f"INSERT INTO `{self.table_name}` " "(no_str, league, datetime_, home, away, a_sf, h_sf, " "battle, a_rate, h_rate, a_ds, h_ds) " "VALUES (%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s,%s)" ) try: self.cursor.execute(sql, ( item["no_str"], item["league"], item["datetime"], item["home"], item["away"], item["a_sf"], item["h_sf"], item["battle"], item["a_rate"], item["h_rate"], item["a_ds"], item["h_ds"], )) self.conn.commit() except Exception as e: logging.error(f"insert failed: {e}") self.conn.rollback() return item建表语句参考:
CREATE TABLE `game_baskstation` ( `id` INT NOT NULL AUTO_INCREMENT, `no_str` VARCHAR(32) DEFAULT NULL, `league` VARCHAR(64) DEFAULT NULL, `datetime_` VARCHAR(32) DEFAULT NULL, `home` VARCHAR(64) DEFAULT NULL, `away` VARCHAR(64) DEFAULT NULL, `a_sf` VARCHAR(32) DEFAULT NULL, `h_sf` VARCHAR(32) DEFAULT NULL, `battle` VARCHAR(128) DEFAULT NULL, `a_rate` VARCHAR(16) DEFAULT NULL, `h_rate` VARCHAR(16) DEFAULT NULL, `a_ds` VARCHAR(32) DEFAULT NULL, `h_ds` VARCHAR(32) DEFAULT NULL, PRIMARY KEY (`id`) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;4. 验证请求:跑起来看 json 结果
配置写完,先别急着落库,用-o参数直接输出 json 验证字段是否拼对:
scrapy crawl bask_station -o output/test.json -s LOG_LEVEL=INFO跑完之后打开output/test.json,正常应该看到类似这样的结构:
[ { "no_str": "周一301", "league": "NBA", "datetime": "01-15 09:30", "home": "湖人", "away": "凯尔特人", "a_sf": "6胜/4负", "h_sf": "7胜/3负", "battle": "主队湖人5胜5负", "a_rate": "60%", "h_rate": "70%", "a_ds": "112/108", "h_ds": "115/105" } ]校验动作分三步:第一,看条数,列表接口返回多少场,json 里就应该有多少条,少了说明详情请求有失败;第二,看字段有没有None或空字符串,尤其是a_ds、h_ds这种拼接字段,如果详情接口返回结构变了,这里会直接报 KeyError;第三,看中文是否正常,如果全是\u开头,检查FEEDS里的ensure_ascii是不是漏了。
如果想把 json 结果丢给模型做字段语义校验,可以用 TaoToken 的模型对话页 https://taotoken.net/models?utm_source=taotoken_aicg_blog_end&utm_content=models&utm_campaign=rewrite ,把一段 json 贴进去问「a_ds 和 h_ds 分别代表什么」,比翻文档快。接入文档在 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite ,里面有 API 调用的完整参数说明。
5. 本篇常见错排查
5.1 403 Forbidden 或返回空 json
最常见的原因是请求头不全。竞彩网的接口对Referer和User-Agent比较敏感,settings.py里DEFAULT_REQUEST_HEADERS必须补上这两个。如果还是 403,检查ROBOTSTXT_OBEY是不是True,改成False。
5.2 callback 报 TypeError
TypeError: 'NoneType' object is not callable或者
TypeError: data_analysis() missing 1 required positional argument基本都是callback写成了callback=self.data_analysis(),把括号去掉就行。callback传的是函数对象,不是调用结果。
5.3 allowed_domains 拦截详情请求
列表接口和详情接口虽然都在webapi.sporttery.cn下,但如果你后面换了别的域名,比如webapi.sporttery.cn之外的接口,allowed_domains里没加,Scrapy 会直接过滤掉请求,日志里显示Filtered offsite request。解决办法是把新域名加进allowed_domains,或者临时用dont_filter=True。
5.4 json 输出中文乱码
FEEDS里没写ensure_ascii: False,或者写成了ensure_ascii: false(Python 里布尔值首字母大写)。正确写法是"ensure_ascii": False。
5.5 详情接口 KeyError
详情接口返回的 json 结构如果变了,比如scoreAvg下面字段名改了,value["scoreAvg"]["awayGoalAvgCnt"]就会抛 KeyError。排查方法是先把response.text打印出来,看实际结构,再对照代码改字段名。建议在data_analysis里加一层try/except,把失败的matchId记下来,方便回头补抓。
def data_analysis(self, response: TextResponse): item = response.meta["key"] try: value = json.loads(response.text)["value"] # ... 字段拼装 yield item except KeyError as e: self.logger.error(f"field missing: {e}, url={response.url}")5.6 数据库连接报字符集错误
pymysql 连接时charset不能写成utf-8,要写成utf8mb4。端口号必须是整型,写成字符串会报TypeError。
6. 长期跑数据,Key 和配置怎么管
如果你只是偶尔跑一次竞彩 NBA 受注数据,上面这套配置够用了。但如果要做成每天定时抓、多联赛并行、抓完还要做字段校验和异常告警,那 Key 和配置的管理就要提前规划。我的做法是把 TaoToken 的 Key、数据库连接、目标联赛 id 全部收进config.py,爬虫代码里只引用变量,不写死任何敏感信息。这样换环境的时候只改一个文件。
长期做编码和 Agent 任务的话,Coding Plan 页面 https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding-plan&utm_campaign=rewrite 有对应的方案说明,适合把模型调用嵌进日常开发流程。ClaudeCode 相关的接入说明在 https://taotoken.net/claudecode-anthropic?utm_source=taotoken_aicg_blog_end&utm_content=claudecode-anthropic&utm_campaign=rewrite ,如果你用 ClaudeCode 做爬虫脚本的辅助编写,可以参考。
最后提醒一句:抓数据要控制频率,DOWNLOAD_DELAY别设太小,CONCURRENT_REQUESTS也别开太高,竞彩网这类接口对高频请求是有风控的。跑之前先小批量验证,确认字段和条数都对,再放开跑全量。