图像与视频理解
识别图片视频里有什么:分类、检测、分割、OCR
数据截至 8/1 18:48(构建快照,正在获取最新)
按仍在维护排序:先筛掉近 30 天没有提交的项目,再按热度排。 高 star 但早已停更的项目不会出现在这里。
- 1
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
- 2
Tesseract Open Source OCR Engine (main repository)
- 3
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
- 4
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
- 5
We write your reusable computer vision tools. 💜
- 6
A community-supported supercharged document management system: scan, index and archive all your documents
- 7
ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images, text, and various file types to a wide range of destinations.
- 8
NVR with realtime local object detection for IP cameras
- 9
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
- 10
Label Studio is a multi-type data labeling and annotation tool with standardized output format
- 11
Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs.
- 12
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
- 13
Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection
- 14
一个简洁优雅的词典翻译 macOS App。开箱即用,支持离线 OCR 识别,支持有道词典,🍎 苹果系统词典,🍎 苹果系统翻译,OpenAI,Gemini,DeepL,Google,Bing,腾讯,百度,阿里,小牛,彩云和火山翻译。A concise and elegant Dictionary and Translator macOS App for looking up words and translating text.
- 15
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
- 16
视觉小说翻译器 / Visual Novel Translator
- 17
A fast, helpful, and open-source document parser
- 18
Refine high-quality datasets and visual AI models
- 19
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
- 20
Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)
- 21
Techniques for deep learning with satellite & aerial imagery
- 22
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
- 23
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM 3, and Qwen3-VL.
- 24
RF-DETR is a real-time object detection and segmentation model architecture developed by Roboflow, SOTA on COCO, designed for fine-tuning. [ICLR 2026]
- 25
A ready-to-go translation ocr tool developed with WPF/WPF 开发的一款即用即走的翻译、OCR工具
- 26
📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.
- 27
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
- 28
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
另有 54 个项目因近期无提交或缺少数据未列入。
图像与视频理解要落到业务里,还差什么?
开源项目给的是能力,不是方案。数据怎么接、权限怎么管、上线后谁维护,这些才是落地的真正成本。