拓十年匠心定制 · 商业建站与技术教学双线并行 咨询热线:400-886-1026 service@lmnt.cn
ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

企业微信API接口如何实现文件自动处理?从接收文件到业务解析的完整流程

企业微信API接口如何实现文件自动处理?从接收文件到业务解析的完整流程

最近做的企微二开,客户经常发文件——订单表、合同、发票。之前机器人收到文件只会回"已收到"。业务方要求文件能自动处理:识别是什么文件、提取内容、走对应业务流程。这篇重点讲文件处理每个环节调了哪些企微接口——文件消息接收、CDN 下载、消息回复。

底层用的是Eyun 平台开放的企微 API,统一 POST+JSON,鉴权用 App Token 加 appid,响应封套{code, data, detail, message, time},code 为 0 成功。

文件消息接收:Webhook 拿 fileId

客户发文件,Webhook 回调里消息类型是file,带fileId、fileName、fileSize:

@app.route("/wx-api/webhook/", methods=["POST"]) def webhook(): payload = request.json["data"] if payload["msgType"] != "file": return "ok" file_id = payload["fileId"] file_name = payload["fileName"] appid = payload["appid"] from_uin = payload["fromUin"] # 异步处理文件 process_file.delay(appid, from_uin, file_id, file_name) return "ok"

文件处理耗时长,不能在回调里同步做,入队异步处理,先回"收到文件,处理中"。

文件下载:调 CDN 接口

拿到 fileId 后调 CDN 模块下载文件:

def download_file(appid, file_id, save_path): resp = requests.post( f"{BASE}/wx-api/api/cdn/download", headers=HEADERS, json={"appid": appid, "fileId": file_id}, stream=True ) if resp.json().get("code") == 0: with open(save_path, "wb") as f: for chunk in resp.iter_content(8192): f.write(chunk) return save_path raise Exception(resp.json()["message"])

下载是流式的,大文件不会爆内存。下载完存 OSS,本地不留。CDN 接口的 fileId 是回调里给的,不能自己拼。接口参数在Eyun 开发文档。

文件类型识别:按扩展名和内容

下载完识别文件类型,决定怎么解析:

import os def detect_file_type(file_name, file_path): ext = os.path.splitext(file_name)[1].lower() if ext in [".xlsx", ".xls", ".csv"]: return "excel" elif ext in [".pdf", ".doc", ".docx"]: return "document" elif ext in [".jpg", ".png", ".jpeg"]: return "image" else: return "other"

类型识别按扩展名加内容双重判断,防止客户改扩展名骗识别。识别错类型,解析全错。

Excel 解析:提取订单数据

Excel 用 openpyxl 读,提取结构化数据。客户的表头可能不统一,做别名匹配:

import openpyxl def parse_excel(file_path): wb = openpyxl.load_workbook(file_path) ws = wb.active headers = [c.value for c in ws[1]] # 表头别名映射 col_map = {} for i, h in enumerate(headers): if h in ["订单号", "订单编号", "order_no"]: col_map["order_no"] = i elif h in ["产品", "商品名称", "product"]: col_map["product"] = i elif h in ["数量", "qty"]: col_map["quantity"] = i orders = [] for row in ws.iter_rows(min_row=2, values_only=True): order = {k: row[v] for k, v in col_map.items() if v < len(row)} if order.get("order_no"): orders.append(order) return orders

解析完校验——订单号不为空、数量是数字。校验不通过的行标记错误。

业务回写:进订单系统

解析成功的订单批量创建。回写要幂等——同一文件重复上传不重复创建。幂等键用文件名加内容哈希:

import hashlib def create_orders(orders, file_name, file_content): idem_key = f"file:{hashlib.md5((file_name+file_content).encode()).hexdigest()}" if redis.set(idem_key, "1", nx=True, ex=86400): return order_api.batch_create(orders) else: return {"status": "duplicate"}

不做幂等,客户重发一次文件,订单建两遍。

处理结果回复:调消息接口

处理完调消息接口把结果回给客户:

def reply_result(appid, to_uin, total, success, failed_rows): text = f"文件处理完成:共 {total} 行,成功 {success} 行" if failed_rows: text += f",失败 {len(failed_rows)} 行:\n" for r in failed_rows[:5]: text += f"- 第{r['row']}行:{r['reason']}\n" requests.post( f"{BASE}/wx-api/api/message/sendText", headers=HEADERS, json={"appid": appid, "to": to_uin, "content": text} )

部分成功要明确告诉客户哪些成功哪些失败,不能含糊说"处理完成"。

异步队列:大文件不阻塞

文件解析耗时长,走异步队列。回调里只下载和入队,worker 处理:

from celery import Celery celery = Celery("file_processor") @celery.task def process_file(appid, from_uin, file_id, file_name): local_path = download_file(appid, file_id, "/tmp") ftype = detect_file_type(file_name, local_path) if ftype == "excel": orders = parse_excel(local_path) result = create_orders(orders, file_name, open(local_path).read()) reply_result(appid, from_uin, len(orders), result["success"], result["failed"]) else: reply_result(appid, from_uin, 0, 0, [{"row": 0, "reason": "不支持的文件类型"}]) os.remove(local_path)

并发控制——同时处理太多文件,解析服务扛不住,加令牌桶限流。

写在最后

文件自动处理这套东西,本质是用好企微的文件和消息接口——Webhook 收文件消息拿 fileId、CDN download 下载文件、sendText 回处理结果。文件解析(类型识别、Excel 提取、业务回写)在应用层做。接口路径、参数、CDN 下载方式在Eyun 开发文档里。把接口调对、解析做扎实,客户发的文件真正能被系统处理——而不是躺在消息里等人手动导。凭证和接入地址在Eyun 企业微信 API 平台开通。

返回列表