Vision & Multimodal plugins

DSH vision plugins add image understanding, OCR, visual grounding, screenshot analysis, and multimodal workflows to DeepSeek Harness. They are how a text-only model reads an interface, a diagram, or a screenshot.

3 reviewed · 126 cataloged

Recommended projects

Cataloged projects

Repositories with traceable evidence of a DeepSeek Harness relationship. They have not been reviewed, install tested, or security checked — the signal below each one is the whole of what the registry currently knows.

How cataloging works
Cataloged

修改DSH的背景,支持静态动态背景,支持网页图片视频,支持修改透明度

README documents a DSH install command Package metadata declares a DSH dependency

dsh-computer-use

ThreeBody6666
Cataloged

Native Windows Computer Use and configurable vision tools for DeepSeek Harness.

README documents a DSH install command Package metadata declares a DSH dependency

dsh-vision-plugin

Xin-Zhang-IceMan
Cataloged

DeepSeek Harness 视觉插件:让纯文本模型拥有视觉能力 / Vision plugin for DSH: vision_analyze tool + automatic image transcription for text-only models.

README documents a DSH install command Package metadata declares a DSH dependency

dsh-eyes

qing9835
Cataloged

DSH 视觉模型插件:为无视觉能力的文本模型提供图片识别。粘贴/拖入/导入的图片自动交给 OpenAI 兼容视觉模型识别为文字并发送进对话,支持多轮复核(vision_ask)、多提供商配置。

README documents a DSH install command Package metadata declares a DSH dependency
Cataloged

KoboldCpp for DeepSeek Harness - a tool plugin that lets the harness online model hand repetitive text and vision (OCR) labor to a local KoboldCpp (llama.cpp) server.

Package metadata declares a DSH dependency

dsh-vision-bridge

TwistedRiCen
Cataloged

DSH-native Vision Evidence bridge for text-only reasoning models with native image attachments and strict multi-image validation.

README documents a DSH install command Package metadata declares a DSH dependency

Vision & Multimodal plugins: common questions

What is a DSH vision plugin?
A plugin that turns images into something the model can reason about — extracted text, layout structure, or a described scene.
Can DeepSeek Harness process images without a plugin?
Only when the selected model is itself multimodal. A vision plugin is what lets a text-only model work from images.
How do vision plugins handle my images?
It varies. Some run entirely locally, others send images to a hosted service. Each record lists the declared network permissions.