拓十年匠心定制 · 商业建站与技术教学双线并行 咨询热线:400-886-1026 service@lmnt.cn
ARTICLE DETAIL

资讯详情

深耕网站建设与运营推广的一线实战洞察。

OpenGlass 视觉模型评测实录:基于模糊抽象样本的 VLM 图像描述能力对比(img_20 案例)

OpenGlass 视觉模型评测实录:基于模糊抽象样本的 VLM 图像描述能力对比(img_20 案例)
  • 人工智能
  • AI 应用
  • 智能硬件
  • 本地部署
  • 可穿戴
  • AI Agent

【免费下载链接】OpenGlass

Turn any glasses into AI-powered smart glasses

项目地址:https://gitcode.com/GitHub_Trending/op/OpenGlass
点击查看免费下载

导读

本文以 OpenGlass 智能眼镜开源项目(Turn any glasses into AI-powered smart glasses)中的一张真实评测样本prompts/series_1/img_20.jpeg为对象,完整复盘该仓库的多模态视觉语言模型(VLM)图像描述流水线是如何对同一张"低辨识度模糊抽象图"进行多模型对比评测的。你将看到:评测脚本如何批量驱动 Ollama 上 4 个模型(默认 moondream、llava-llama3、llava:34b-v1.6、moondream:1.8b-v2-fp16)生成描述、描述文本如何在img_20.md中按固定格式落盘,以及四个模型针对这张无实体、仅由绿/棕模糊渐变与噪点构成的图片分别给出了哪些差异巨大的输出——并由此理解"图像描述质量评价(含幻觉、过度泛化、细节漏报)"这一环节在 OpenGlass Agent 视觉问答链路中的位置与作用。

一、样本背景:img_20 到底是什么样的图片

先对齐证据。prompts/series_1/目录中存放了 57 张评测图(img_1 至 img_57,jpeg 与同名 md 成对出现),img_20.jpeg是其中一张竖幅图(600×800)。从画面本身看,它没有任何可命名的实体对象:画面主体是左半侧暗棕/土黄调与右半侧青绿调之间柔和渐变融合的模糊色块,底部融为暗青黑,伴随明显的颗粒噪声,全图没有清晰轮廓、文字或纹理。

关键点:这不是"渲染失败的废图",而是被刻意纳入评测集的语义模糊级(low-recognition / abstract)样本。它在评测中的作用是专门探测 VLM 面对"不存在明确语义目标"的输入时,是会老老实实描述色彩与画质,还是会脑补出不存在的物体。

二、评测机制:prompts/generate.ts 如何批量驱动多模型

img_20.md不是手写的,而是由仓库根目录的 prompts/generate.ts 批量脚本自动生成的。其流程如下:

  1. 遍历prompts/下的每个 series 目录,收集所有.jpeg文件(见 generate.ts),并把输出路径定为同名.md文件;
  2. 依次执行 4 组"模型描述"测试,每组把模型输出以####标题####分隔块追加进outputs:
    • Description:调用 imageDescription 的默认模型moondream:1.8b-v2-fp16;
    • Description (llava-llama3):传入llava-llama3;
    • Description (llava:34b-v1.6):传入llava:34b-v1.6;
    • Description (moondream:1.8b-v2-fp16):显式传入moondream:1.8b-v2-fp16,与默认模型结果互为对照;
  3. 脚本末尾用fs.writeFileSync把拼接结果写回对应 md 文件(见 generate.ts),并使用cli-progress打印进度条。

这意味着img_20.md中每一个####Description (xxx)####块,都对应一次真实的 Ollama 推理请求。

三、底层调用链:从 generate.ts 到 Ollama 推理

imageDescription的核心实现位于 imageDescription.ts:

export async function imageDescription(src: Uint8Array, model: KnownModel = 'moondream:1.8b-v2-fp16'): Promise<string> { return ollamaInference({ model: model, messages: [{ role: 'system', content: 'You are a very advanced model and your task is to describe the image as precisely as possible. Transcribe any text you see.' }, { role: 'user', content: 'Describe the scene', images: [src], }] }); }

随后进入 ollama.ts 的ollamaInference:

  • 请求体:向keys.ollama(即环境变量EXPO_PUBLIC_OLLAMA_API_URL,见 keys.ts)POSTstream: false的/api/chat请求,携带model与 messages;
  • 图片编码:Uint8Array图片经 toBase64 转为 base64 字符串后放入images字段;
  • 文本规整:消息与输出都经过 trimIdent 去除公共缩进,避免 prompt 里的多行缩进污染模型输入;
  • 容错重试:请求被backoff包裹(实现见 time.ts 的createBackoff),失败时按指数退避重试(默认 minDelay 250ms、maxDelay 1000ms、maxFailureCount 50),保证长批量评测的稳定性。

支持的全部模型在KnownModel联合类型中定义(见 ollama.ts),包括llama3、llama3-gradient、llama3:8b-instruct-fp16、llava-llama3、llava:34b-v1.6、moondream:1.8b-v2-fp16等。

四、四模型实测输出:同一张模糊图,四种截然不同的"解读"

以下为img_20.md中记录的、四组真实模型输出的完整对比(保持原文英文,随后给出分析):

评测分组模型输出原文
Descriptionmoondream:1.8b-v2-fp16(默认)The image features a green and gray color scheme, with the majority of the scene being in shades of green. There is a large section of green that appears to be blurry or fuzzy, giving it an artistic or dreamy appearance. The background consists of two vertical lines, one on each side of the frame, which further emphasize the green elements within the image.
Description (llava-llama3)llava-llama3The image captures a close-up view of an abstract design on a dark blue background. The design is composed of small, square-shaped elements that are colored in shades of green and orange. These elements are arranged in a diagonal line across the image, creating a sense of movement and dynamism. The color gradient varies from green to orange, adding depth and dimension to the design. The overall effect is one of modern art, characterized by its use of geometric shapes and abstract colors.
Description (llava:34b-v1.6)llava:34b-v1.6The image you've provided appears to be extremely blurry and low resolution, making it difficult to discern any specific details. The colors are muted, with shades of green and blue dominating the scene. Due to the lack of clarity, I am unable to provide a detailed description or transcribe text if there is any. If you have a higher-quality image or more context, please share so that I can assist you better.
Description (moondream:1.8b-v2-fp16)moondream:1.8b-v2-fp16The image features a beautiful view of the sky, with its blue and green hues blending together to create an awe-inspiring scene. The sky appears to be slightly blurred, adding a softness and dreamy quality to the overall composition.

五、逐条点评:从这四条输出你能学到什么

5.1 默认 moondream:贴近画面,但引入"无中生有"的线条

moondream(默认模型)正确地抓住了"green and gray / blurry or fuzzy / artistic or dreamy"这些画质与色调事实,但对画面右下方实际的噪点/渐变过渡物,它描述为"two vertical lines, one on each side of the frame"——这与样本实际内容并不完全吻合,属于低置信度幻觉:模型试图把不存在的结构"解释"成线条。这正是评测模糊样本的第一个价值:观察模型是否会虚构结构。

5.2 llava-llama3:最严重的过度泛化

llava-llama3 的输出把模糊色块脑补成了一幅"由绿色与橙色小方块构成、斜向排列、具现代艺术感"的抽象设计。原文画面并无离散的方块元素,橙色也并非画面主色。这说明 llava-llama3 在低信息输入下用训练先验补全了视觉内容——它倾向于把"模糊"翻译成"艺术"来维持描述的连贯性。对 OpenGlass 这类实时眼镜场景而言,这种过度泛化会直接污染后续的问答证据。

5.3 llava:34b-v1.6:最诚实的"拒绝回答"

llava:34b-v1.6 是四个输出中唯一承认自身局限的:它明确说出"extremely blurry and low resolution……unable to provide a detailed description",并提示用户提供更高质量图片。它正确报告了 muted colors、green and blue 等有限信息,但把右侧青绿区误报为蓝色。从评测角度,这代表"低信息输入下的保守策略"——宁可少说也不编造。

5.4 默认 moondream(第二组):与 5.1 同一模型的随机性对比

第二组 moondream 输出把它描述为"a beautiful view of the sky……blue and green hues blending"——与第一组 moondream 输出并不一致。同一模型、同一图片、两次推理产生不同描述(一个说绿色/灰色+两条线,一个说天空/蓝绿渐变),说明模糊样本下的输出具有明显的采样随机性,也解释了评测脚本为何用两个分组重复调用默认模型。

5.5 四模型对比小结

维度默认 moondreamllava-llama3llava:34b-v1.6moondream(重复组)
色调还原绿/灰(基本准确)绿/橙/深蓝背景(部分虚构)绿/蓝(蓝为误报)蓝/绿(蓝为误报)
结构描述虚构两条竖线虚构方块+斜向排列承认无法分辨虚构"天空"
画质判断模糊/梦幻现代艺术模糊/低分辨率(准确)模糊/柔和
幻觉程度中高低(保守)中

六、这套评测在 OpenGlass Agent 中的实际用途

6.1 图像描述是"先描述、后问答"管线的第一环

在 Agent.ts 中,用户拍下的照片会先经过imageDescription生成文本描述并存入#photos(见 Agent.ts),当用户提问时,answer()把所有照片描述拼接后交给llamaFind(即 Groq 的 llama3-70b-8192,见 imageDescription.ts 与 groq-llama3.ts)回答问题,期间还会调用 OpenAI TTS(见 openai.ts 的textToSpeech)。

因此img_20这类模糊样本一旦被加入评测集,就直接回答了一个工程问题:当用户在低光照/低画质下拍照提问时,哪一层的描述会失真?上文的对比表明,对模糊图最稳妥的策略是 llava:34b-v1.6 的"保守拒绝"或 moondream 的"画质描述",而 llava-llama3 的过度泛化会作为错误证据传给下游问答模型。

6.2 与 imageBlurry 质量闸门的关系

仓库还提供了专门的图像质量判定函数 imageBlurry,用moondream:1.8b-v2-moondream2-text-model-f16回答"这张图是否模糊/损坏/低质量(YES or NO)"。generate.ts 中这段评测被注释掉,说明它属于预留的评测能力:未来可以在描述之前先用 Blurry 闸门拦截低质量照片,避免把模糊图的幻觉描述送入问答管线。img_20 恰好就是这类闸门最典型的拦截对象。

七、如何复现与扩展这套评测

复现img_20.md的生成过程:

  1. 本地安装并启动 Ollama,确保默认模型可用(README 中给出的步骤是ollama pull moondream:1.8b-v2-fp16,见 README.md);
  2. 在 keys.ts 中通过环境变量EXPO_PUBLIC_OLLAMA_API_URL配置 Ollama 服务地址(默认应为http://localhost:11434/api/chat,见 README.md);
  3. 在prompts/series_1/下放入新的.jpeg评测图,运行ts-node prompts/generate.ts(或按 package.json 中的脚本),即可得到带 4 个####Description####分块的同名.md文件;
  4. 若需加入更多模型,在KnownModel(ollama.ts)中扩展模型名,并在 generate.ts 追加一组runTest调用即可;取消 generate.ts 的注释还可开启 Blurry 质量判定评测。

八、结论

prompts/series_1/img_20.md这份评测记录展示了 OpenGlass 团队评估多模态模型的一种务实方法:用真实眼镜场景可能遭遇的低质量照片,对比多个开源 VLM 的描述输出,从而为下游 Agent 问答管线选择更稳的描述策略。从本文的逐条对比可以看到:模糊抽象样本能最有效地暴露模型的幻觉与过度泛化倾向,而"保守描述 + 质量闸门拦截"是当前代码库(imageDescription.ts、imageBlurry.ts、Agent.ts)给出的可行工程答案。如果你正在为自己的视觉 Agent 选择图像描述模型,建议用类似 img_20 的样本集先跑一遍多模型对比,再决定默认模型与质量过滤策略。

  • 人工智能
  • AI 应用
  • 智能硬件
  • 本地部署
  • 可穿戴
  • AI Agent

【免费下载链接】OpenGlass

Turn any glasses into AI-powered smart glasses

项目地址:https://gitcode.com/GitHub_Trending/op/OpenGlass
点击查看免费下载

相关推荐

上一篇:终极Google Breakpad符号文件生成指南:从调试信息到可读堆栈跟踪的完整流程
下一篇:深度解析:Readium-js-viewer的架构设计与模块化实现原理

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

返回列表