大模型底座 · 中文选型解读

gemma-4-26B-A4B-it

Hugging Faceapache-2.0

这是谷歌DeepMind推出的Gemma 4系列26B参数指令微调多模态大模型,支持图文输入生成文本,可适配不同算力的部署场景。

12.2M次下载近 30 天仍在维护维护状态apache-2.0 · 可评估商用商用提醒
在 Hugging Face 查看官方项目
适合解决作为问答、生成和智能体应用的基础模型
更适合正在比较模型能力、成本与部署方式的团队
投入判断上手门槛:较高。需要评测真实业务数据与许可边界
一分钟看懂

这个项目值得继续研究吗?

AI 依据上游资料解读 · 2026/8/3

这是谷歌DeepMind推出的Gemma 4系列26B参数指令微调多模态大模型,支持图文输入生成文本,可适配不同算力的部署场景。

解决什么问题
解决企业搭建多模态AI应用时,闭源模型成本高、权限不灵活,开源模型多模态处理能力不足、适配场景有限,难以平衡效果与部署成本的痛点。
适合什么团队
适合需要开发多模态问答、智能内容生成、智能客服等应用,具备基础大模型部署能力的企业业务及技术团队使用。
使用前注意
本版本仅支持图文输入输出,如需音视频能力需选择同系列更小参数版本;采用Apache 2.0许可可商用,26B参数需匹配对应GPU算力部署。

本页用于缩短初步筛选时间,不构成技术、采购或法律结论。 正式使用前请在真实业务数据上验证,并以官方说明与许可证为准。

可核对的事实层

官方资料与来源

查看来源 →
  • transformers
  • safetensors
  • gemma4
  • image-text-to-text
  • conversational
  • eval-results
  • endpoints_compatible
查看上游原始说明节选

任务类型:image-text-to-text

<div align="center">
  <img src=https://ai.google.dev/gemma/images/gemma4_banner.png>
</div>


<p align="center">
    <a href="https://huggingface.co/collections/google/gemma-4" target="_blank">Hugging Face</a> |
    <a href="https://github.com/google-gemma" target="_blank">GitHub</a> |
    <a href="https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/" target="_blank">Launch Blog</a> |
    <a href="https://ai.google.dev/gemma/docs/core" target="_blank">Documentation</a> |
    <a href="https://arxiv.org/abs/2607.02770" target="_blank">Technical Report</a>
    <br>
    <b>License</b>: <a href="https://ai.google.dev/gemma/docs/gemma_4_license" target="_blank">Apache 2.0</a> | <b>Authors</b>: <a href="https://deepmind.google/models/gemma/" target="_blank">Google DeepMind</a>
</p>

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. 

Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: **E2B**, **E4B**, **12B**, **26B A4B**, and **31B**. Their diverse sizes make them deployable in environments ranging from high-end phones to laptops and servers, democratizing access to state-of-the-art AI.

Gemma 4 introduces key **capability and architectural advancements**:

* **Reasoning** – All models in the family are designed as highly capable reasoners, with configurable thinking modes.

* **Extended Multimodalities** – Processes Text, Image with variable aspect ratio and resolution support (all models), Video, and Audio (featured natively on the E2B, E4B, and 12B models).

* **Diverse & Efficient Architectures** – Offers Dense and Mixture-of-Experts (MoE) variants of different sizes for scalable deployment.

* **Optimized for On-Device** – Smaller models are specifically designed for efficient local execution on laptops and mobile devices.

* **Increased Context Window** – The small models feature a 128K context window, while the medium models support 256K.

* **Enhanced Coding & Agentic Capabilities** – Achieves notable improvements in coding benchmarks alongside native function-calling support, powering highly capable autonomous agents.

* **Native System Prompt Support** – Gemma 4 introduces native support for the `system` role, enabling more structured and controllable conversations.

## **Models Overview**

Gemma 4 models are designed to deliver frontier-level performance at each size, targeting deployment scenarios from mobile and edge devices (E2B, E4B) to consumer GPUs and workstations (12B, 26B A4B, 31B). They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.

The models employ a hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global. This hybrid design delivers the processing speed and low memory footprint of a lightweight model without sacrificing the deep awareness required for complex, long-context tasks. To optimize memory for long contexts, global layers feature unified Keys and Values, and apply Proportional RoPE (p-RoPE). 

### Dense Models

| Property | E2B | E4B | 12B Unified | 31B Dense |
| :---- | :---- | :---- | :---- | :---- |
| **Total Parameters** | 2.3B effective <br> (5.1B with embeddings) | 4.5B effective <br> (8B with embeddings) | 11.95B | 30.7B |
| **Layers** | 35 | 42 | 48 | 60 |
| **Sliding Window** | 512 tokens | 512 tokens | 1024 tokens | 1024 tokens |
| **Context Length** | 128K tokens | 128K tokens | 256K tokens  | 256K tokens  |
| **Vocabulary Size** | 262K | 262K | 262K | 262K 

上游文档较长,此处为节选。完整内容见官方项目。