- 人工智能
- 深度学习
- 推理引擎
- 本地部署
- 嵌入式
- 物联网
【免费下载链接】tflite-micro
Infrastructure to enable deployment of ML models to low-power resource-constrained embedded targets (including microcontrollers and digital signal processors).
本文以 TensorFlow Lite Micro(TFLM)官方内存管理文档为核心,结合本仓库源码,完整讲解 TFLM 在单片机上管理共享 Tensor Arena 的"在线"分配策略:Head(非持久化)、Temporary(临时)、Tail(持久化)三段式布局、离线内存规划(Offline Memory Plan)的 flatbuffer 编码格式,以及通过RecordingMicroInterpreter审计各分配分区的实战方法。读完本文,你将能够自主计算模型所需的 arena 大小、读懂PrintAllocations()的输出、并借助录制式 API 定位内存瓶颈,为资源受限的嵌入式目标裁剪内存占用。
总览:单一缓冲区的三段式分配
TFLM 在默认 API 下采用"在线"(online)分配策略:模型推理所需的全部"工作空间"都被放进一个char或int8_t数组(称为 Tensor Arena / tensor arena)。这个缓冲区有两种注入方式:
- 直接传入
tflite::MicroInterpreter构造函数; - 先构造一个
tflite::MicroAllocator实例,再把它传入tflite::MicroInterpreter构造函数(后者也用于录制分配、或在多个解释器间共享分配器,见 micro_interpreter.h)。
在内部,tflite::MicroAllocator把所有分配划分为三个区域:
- Head(头部):非持久化分配;
- Temporary(临时区):短生命周期的"作用域"分配;
- Tail(尾部):持久化分配。
典型布局如下(低地址在左,高地址在右):
-------------------------------------------------------------------------------- | | | | | HEAD |<-- TEMPORARY -->| TAIL | | | | | -------------------------------------------------------------------------------- * Lowest Address Highest Address *从源码结构看,这种"头尾相向生长"的设计对应 single_arena_buffer_allocator.h 中的SingleArenaBufferAllocator:它同时继承INonPersistentBufferAllocator与IPersistentBufferAllocator,头侧从低地址向上分配(head watermark),尾侧从高地址向下分配,两个方向最终在 arena 中间会合,中间未使用的空隙就是临时区。
Head Section:非持久化共享张量区
Head 区存放的是共享的 Tensor 缓冲区(各算子激活张量的数据)。这一区不会做零碎的小块迭代分配,而是由规划器一次性确定整个区段的长度。
该区段的长度由tflite::GreedyMemoryPlanner管理。GreedyMemoryPlanner会观察整个模型的计算图,尽可能多地复用缓冲区,从而把 Head 的长度压到最小。它的贪心算法在 greedy_memory_planner.h 的注释中有完整描述:
- 客户端通过
AddBuffer()录入每个缓冲区的信息(大小、首次使用时间、最后使用时间); - 当需要
GetOffsetForBuffer()时,若当前规划不是最新的,则触发CalculateOffsetsIfNeeded()重新计算; - 缓冲区按大小降序排列,最大的缓冲区放在偏移 0 处;
- 其余缓冲区按大小降序遍历,找出所有与它在时间上同时存活的缓冲区,把当前缓冲区塞进第一个能容纳它的空隙;
- 若没有足够大的空隙,则放到最后一个同时活跃缓冲区的后面;
- 直到所有缓冲区都放置完毕。
需要说明的是,该算法并不保证最优放置(这本质上是 NP-Complete 问题),但在实践中已能给出相当不错的布局。此外,GreedyMemoryPlanner::preserves_all_tensors()返回false——推理过程中不再活跃的张量数据会被覆盖复用,这是 TFLM 省内存的关键手段(详见 greedy_memory_planner.h)。
Head 区的 Tensor 缓冲区可以通过tflite::MicroInterpreter上的TfLiteEvalTensor或TfLiteTensor实例访问。使用规划器时需要注意它的 scratch 开销:规划器自身需要一个 scratch 缓冲区,每规划一个缓冲区约占用 36 字节(per_buffer_size(),见 greedy_memory_planner.h),scratch 内存不足时AddBuffer()会报错。
离线规划的 Tensor 分配(Offline Planned Tensor Allocations)
全部或部分 Tensor 也可以采用离线规划器(offline planner)来分配。离线规划器在宿主机(如 PC)上执行 Tensor 布局计算,把结果作为模型元数据写入 flatbuffer,TFLM 运行时直接读取该方案,从而获得比在线贪心算法更紧凑的布局。
对 subgraph 的tensors:[Tensor]列表中每一个非常量Tensor,离线方案给出它相对 arena Head 区起始位置的字节偏移;偏移为-1表示该 Tensor 在运行时由tflite::GreedyMemoryPlanner现场分配。只要离线规划器确认两个缓冲区的数据不会同时使用,就允许它们重叠。注意,这一语义与GreedyMemoryPlanner中的kOnlinePlannedBuffer = -1常量(见 greedy_memory_planner.h)对应。
离线 Tensor 分配方案编码在模型的metadata:[Metadata]字段中,编码方式如下:
| Metadata 组件 | 值 | |-|-| | name:string | "OfflineMemoryAllocation" | | buffer:unit | 包含离线 Tensor 分配数据的 buffer 索引 |
该 buffer 的内容是一串 32 位整数,格式如下:
| 偏移 | 值 | |-|-| | 0 | 离线分配格式版本号 | | 1 | subgraph 数量 | | 2 | 后续偏移数量:n | | 3 | tensor #0 的字节偏移,或 -1(表示运行时分配) | | 4 | tensor #1 的字节偏移,或 -1(表示运行时分配) | | ... | ... | | 3+(n-1) | tensor #(n-1) 的字节偏移,或 -1(表示运行时分配) |
官方文档特别提醒:偏移 0(版本号)和偏移 1(subgraph 数量)目前会被 micro 内存分配器忽略。在多 subgraph 的情况下,分配器假设所有 subgraph 的 Tensor 按顺序首尾拼接:第一个 subgraph 的全部 Tensor 在前,接着是第二个 subgraph 的 Tensor,依此类推。
tflite::GreedyMemoryPlanner会把这份离线方案当作固定的相对 Head 区起始位置的偏移,然后在那些固定偏移周围,尝试为其他 Tensor(例如运行时通过TfLiteContext的RequestScratchBufferInArenaAPI 添加的 scratch Tensor)寻找合适的位置。
基于 C++ 结构体的替代方案:NonPersistentMemoryPlannerShim
除 flatbuffer 元数据外,仓库还提供了另一种处于实验阶段的离线规划入口:NonPersistentMemoryPlannerShim(见 docs/offline_memory_plan.md 与 non_persistent_buffer_planner_shim.h)。它允许用 C++ 结构体直接指定每个非持久化缓冲区的偏移,且最终二进制中不会包含任何GreedyMemoryPlanner相关符号,从而进一步缩小固件体积:
const struct BufferPlan kOfflineNonPersistentBufferPlan = { .buffer_count = 9, .buffer_plan_entries = { [0] = { .offset = 0 }, [1] = { .offset = 400 }, [2] = { .offset = 801 }, [3] = { .offset = 400 }, [4] = { .offset = 811 }, [5] = { .offset = 601 }, [6] = { .offset = 814 }, [7] = { .offset = 601 }, [8] = { .offset = 801 }, } };然后把规划器实例注入MicroAllocator,再交给解释器:
// The arena includes both persistent buffers and non-persistent buffers. constexpr int kArenaSize = 2*1048; uint8_t tensor_arena[kArenaSize]; tflite::NonPersistentMemoryPlannerShim planner(&kOfflineNonPersistentBufferPlan); tflite::MicroAllocator * allocator = tflite::MicroAllocator::Create( tensor_arena, arena_size, &planner); tflite::MicroInterpreter interpreter(model, op_resolver, allocator);Temporary Section:短生命周期作用域分配
临时区用于分配"作用域"(scoped)或短期、不保证长期有效的缓冲区。它的特点是:
- 分配从当前Head 区段的末尾地址开始,向 Tail 方向生长;
- 一整条分配链可以被重置(且在调整 Head 之前必须先重置),重置后当前分配起始地址回到 Head 区段的末尾。
TFLM 目前用临时区来分配生命周期至少覆盖一次方法调用的大型 C 结构体或 scratch 内存。典型场景见 docs/online_memory_allocation_overview.md:在算子的TfLiteRegistration::prepare()阶段,算子需要通过TfLiteTensor结构体读取张量信息,TFLM 会按需在临时区逐个分配这些较重的结构体;每个算子 prepare 结束后,MicroInterpreter调用MicroAllocator::FinishPrepareNodeAllocations(),重置临时分配并把该算子的 scratch buffer 请求固化进 Head 区。也正因如此,TfLiteTensor只存在于 prepare 阶段,之后只能通过更轻量的TfLiteEvalTensor访问数据。
Tail Section:持久化分配区
Tail 区存放 TFLM 使用的所有持久化分配,例如算子注册信息、TfLiteEvalTensor数组、变量张量缓冲等。该区包含许多大小随机的分配,且向 Head 区段的末尾方向生长——即从高地址向低地址推进。
由于这一区的构成比较杂,TFLM 提供了录制式 API(Recording Memory APIs)来辅助审计该区的内容。源码中对应的持久化分配器实现位于 persistent_arena_buffer_allocator.cc,其测试见 persistent_arena_buffer_allocator_test.cc。
Recording Memory APIs:录制式内存审计
TFLM 提供了一组简单的 API,用于审计共享 Tensor Arena 中的内存使用情况。这些 API按需开启(opt-in),会带来额外的内存开销,并且要求目标平台有可用的调试日志实现(参考实现见 debug_log.cc)。
先看一个最精简的裸机 TFLM 解释器初始化代码:
// Buffer for the tensor arena: size_t tensor_arena_size = 2048; uint8_t tensor_arena[tensor_arena_size]; // Interpreter using the shared tensor arena above: tflite::MicroInterpreter interpreter( tflite::GetModel(my_model_data), ops_resolver, tensor_arena, tensor_arena_size); // Invoke one time which will allocate internals: if (interpreter.Invoke() != kTfLiteOk) { MicroPrintf("Exception during invoke()!"); }要启用录制式审计,只需引入RecordingMicroInterpreter类(头文件 recording_micro_interpreter.h),并把类名从tflite::MicroInterpreter换成tflite::RecordingMicroInterpreter。它继承自MicroInterpreter,内部用RecordingMicroAllocator::Create(...)构造分配器,并在首次Invoke()(或AllocateTensors())后记录全部分配信息。同样调用一次invoke(),再调用PrintAllocations()即可输出详细的分配日志:
// Add an include to the recording API: #include "recording_micro_interpreter.h" // Simply change the class name from 'MicroInterpreter' to 'RecordingMicroInterpreter': tflite::RecordingMicroInterpreter interpreter( tflite::GetModel(my_model_data), ops_resolver, tensor_arena, tensor_arena_size); // Invoke one time which will allocate internals: if (interpreter.Invoke() != kTfLiteOk) { MicroPrintf("Exception during invoke()!"); } // Print out detailed allocation information: interpreter.GetMicroAllocator().PrintAllocations();RecordingMicroInterpreter的头文件注释还给出了一条实用建议:建议把 tensor arena 至少增大 1KB,以确保内部录制有足够的额外内存(见 recording_micro_interpreter.h)。
PrintAllocations()的输出与下面类似(示例输出来自 memory_arena_threshold_test.cc 中的 conv 模型测试):
[RecordingMicroAllocator] Arena allocation total 9568 bytes [RecordingMicroAllocator] Arena allocation head 7744 bytes [RecordingMicroAllocator] Arena allocation tail 1824 bytes [RecordingMicroAllocator] 'TfLiteEvalTensor data' used 360 bytes with alignment overhead (requested 360 bytes for 15 allocations) [RecordingMicroAllocator] 'Persistent TfLiteTensor data' used 0 bytes with alignment overhead (requested 0 bytes for 0 tensors) [RecordingMicroAllocator] 'Persistent TfLiteTensor quantization data' used 0 bytes with alignment overhead (requested 0 bytes for 0 allocations) [RecordingMicroAllocator] 'TfLiteTensor variable buffer data' used 0 bytes with alignment overhead (requested 0 bytes for 0 allocations) [RecordingMicroAllocator] 'NodeAndRegistration struct' used 392 bytes with alignment overhead (requested 392 bytes for 7 NodeAndRegistration structs) [RecordingMicroAllocator] 'Operator runtime data' used 136 bytes with alignment overhead (requested 136 bytes for 5 OpData structs)解读这份输出时要注意:requested bytes是内核实际请求的字节数,used bytes是含对齐(alignment)开销后真实占用的字节数,二者之差即对齐浪费。从上到下,日志先给总览(total/head/tail),再逐项列出各类持久化分配的明细。
从 recording_micro_allocator.h 的源码可以看到,RecordingMicroAllocator内部通过RecordedAllocationType枚举管理录制桶(bucket),每个桶对应一个RecordedAllocation结构体,记录requested_bytes、used_bytes和count(分配次数)。除了上文日志中出现的类型外,枚举中还包含kPersistentBufferData(算子经AllocatePersistentBuffer请求的持久化数据),并在USE_TFLM_COMPRESSION宏开启时额外录制kCompressionData(压缩数据)。这些数据还可以通过GetRecordedAllocation(RecordedAllocationType)按类型编程式读取,用于集成测试或内存回归监控。
仓库中的 memory_arena_threshold_test.cc 正是这种审计能力的典型用例:它用 keyword 模型与 conv 模型分别构造RecordingMicroInterpreter,调用AllocateTensors()后,通过ValidateModelAllocationThresholds()逐一校验 head、tail、各类录制分区的实际字节数,并要求内存增长不超过 3%(kAllocationThreshold = 0.03,见 memory_arena_threshold_test.cc),从而防止 arena 布局无意间膨胀。
各分配分区明细(Allocation Section Details)
PrintAllocations()中每个录制分区的具体含义如下:
'TfLiteEvalTensor data'持有数据类型、维度以及指向 Tensor 缓冲区的指针的 C 结构体数组。它是推理运行时访问张量数据的"轻量级"来源。
'Persistent TfLiteTensor data'比
TfLiteEvalTensor携带更多信息(如量化参数、shape 详情等)的 C 结构体。这个桶只有在通过tflite::MicroInterpreter上的访问器访问张量时才会出现分配记录。'Persistent TfLiteTensor quantization data'分配给持久化
TfLiteTensor结构体的持久化量化数据长度。同样只在通过tflite::MicroInterpreter访问器访问张量时出现。'TfLiteTensor variable buffer data'变量张量(variable tensor)的缓冲数据长度。变量张量的特点是在多次
invoke()调用之间保留数据,例如 LSTM 的状态张量,因此必须持久化存放。'NodeAndRegistration struct'一个同时持有
TfLiteRegistration与TfLiteNode结构体实例的 C 结构体。模型中的每个算子对应一个NodeAndRegistration,其定义见 micro_allocator.h。'Operator runtime data'TFLM 内核缓存的持久化分配数据(例如量化参数、乘数、
OpData结构体等),通常由算子的init()回调返回并挂在TfLiteNode的user_data上。
实践要点与延伸阅读
- arena 需要 16 字节对齐:
MicroAllocator::Create()的注释明确建议使用alignas(16)声明 tensor_arena,否则会浪费一部分头部空间(见 micro_allocator.h)。 - 精确获取最优 arena 大小:
AllocateTensors()之后可以调用MicroInterpreter::arena_used_bytes(),得到实际使用的 arena 字节数,作为裁剪 arena 的依据(见 micro_interpreter.h)。 - 内存回归测试:参考 memory_arena_threshold_test.cc 的模式,把录制式审计与固定阈值结合,纳入 CI 以防止内存布局回退。
- 想深入理解在线分配的完整生命周期(Init / Prepare / Finish 三个阶段),可阅读 docs/online_memory_allocation_overview.md;离线规划的另一种 C++ 结构体方案见 docs/offline_memory_plan.md;
docs/目录下还有 docs/memory_management.md 的配套文档可交叉对照。
- 人工智能
- 深度学习
- 推理引擎
- 本地部署
- 嵌入式
- 物联网
【免费下载链接】tflite-micro
Infrastructure to enable deployment of ML models to low-power resource-constrained embedded targets (including microcontrollers and digital signal processors).
相关推荐
Candle内存管理:Storage与Tensor内存布局优化
Candle内存管理:Storage与Tensor内存布局优化 痛点:为什么需要关注内存管理? 在深度学习框架中,内存管理是性能优化的核心环节。你是否遇到过以下
人工智能大模型机器学习深度学习本地部署模型推理服务CANN社区任务2026
CANN 社区任务 2026 任务介绍 CANN 社区任务 2026 是由 CANN 开源社区发起的年度算子开发任务,面向所有开源贡献者,旨在推动昇腾 AI 生
CANN文档高性能计算Apache Arrow 格式规范解读:Tensor 与 Sparse Tensor 数据结构的 IPC 序列化与内存布局
Apache Arrow 格式规范解读:Tensor 与 Sparse Tensor 数据结构的 IPC 序列化与内存布局 Apache Arrow 的核心价值
数据工程大数据序列化数据分析
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考