AWS Adds Eight Open Models to SageMaker JumpStart

The Core · TL;DR
- AWS added eight open models to SageMaker JumpStart, including FLUX.2-small-decoder, Gemma-4-12B-it, GLM-5.2 FP8, GLM-OCR, LightOnOCR-2-1B, Mellum2-12B-A2.5B-Thinking, NVIDIA-Nemotron-Nano-12B-v2, and langcache-embed-v3-small.
- Standouts include a 1M-token context GLM-5.2 FP8 for agentic engineering, a 9x-smaller LightOnOCR-2-1B claiming SOTA OCR benchmark results, and a Mamba-2/Transformer hybrid Nemotron model with up to 6x inference throughput gains.
- Independent verification could not confirm the exact 'gemma-4-12B-it' and 'FLUX.2-small-decoder' variant names cited by AWS, with external sources instead referencing 'gemma-4-E2B-it' and 'FLUX.2-klein-base-4B'.
- The release continues AWS's push to broaden JumpStart's catalog of compact, efficiency-focused open models for OCR, coding, embeddings, and multimodal reasoning.
Amazon has expanded SageMaker JumpStart with a batch of eight third-party open models spanning image generation, OCR, coding, embeddings, and general-purpose reasoning, giving developers pre-configured access to deploy them directly within AWS infrastructure.
The lineup covers a wide range of use cases. FLUX.2-small-decoder targets faster image generation pipelines, claiming roughly 1.4x quicker decoding at 1.4x lower VRAM usage compared to prior variants. Gemma-4-12B-it, meanwhile, is pitched as a compact multimodal model capable of running on just 16GB of RAM while handling text, image, and audio inputs alongside native function calling for agentic workflows.
On the document-processing side, GLM-OCR and LightOnOCR-2-1B stand out for their small footprints. GLM-OCR, at 0.9B parameters, converts documents including tables and formulas into Markdown, JSON, or LaTeX. LightOnOCR-2-1B is roughly 9x smaller than competing OCR approaches yet reportedly achieves state-of-the-art results on the OlmOCR-Bench benchmark.
Two models are aimed squarely at long-context and engineering-heavy workloads. GLM-5.2 FP8 offers a 1-million-token context window designed for project-scale agentic engineering tasks, while JetBrains' Mellum2-12B-A2.5B-Thinking uses a Mixture-of-Experts design, activating only 2.5B of its 12B total parameters per forward pass across 64 experts (8 active at a time), with a 131,072-token context length.
NVIDIA-Nemotron-Nano-12B-v2 rounds out the reasoning-focused additions with a hybrid Mamba-2 and Transformer architecture, claiming up to 6x higher inference throughput than comparable open models at a 128K context length. Redis contributed langcache-embed-v3-small, an embedding model built for semantic caching in LLM application pipelines.
Naming Discrepancies Flagged
Independent verification of two entries in AWS's announcement ran into inconsistencies. Public records point to a "gemma-4-E2B-it" variant appearing on JumpStart in July 2026, not the "gemma-4-12B-it" named in AWS's January posting, and no external source corroborates a distinct 12B version under that exact name.
Similarly, outside reporting identifies a "FLUX.2-klein-base-4B" model added to JumpStart in May 2026, but no independent source confirms the "FLUX.2-small-decoder" variant AWS describes. The specifications may reflect legitimate model updates or renamed releases, but the exact variant names in AWS's announcement remain unverified against third-party sources.
For teams building on SageMaker, the practical takeaway is straightforward: JumpStart continues to broaden its catalog of smaller, efficiency-oriented open models suited to constrained hardware and specialized tasks like OCR and semantic caching, though buyers should confirm exact model specifications before deployment given the naming ambiguities.
Original reporting and research used to synthesize this article.
- 1GLM-5.2 FP8, NVIDIA-Nemotron-Nano-12B-v2 and GLM-OCR models now available on Amazon SageMaker JumpStartaws.amazon.com
- 2langcache-embed-v3-small, Mellum2-12B-A2.5B-Thinking, and LightOnOCR-2-1B models now available on Amazon SageMaker JumpStartaws.amazon.com
- 3FLUX.2-small-decoder and gemma-4-12B-it models now available on Amazon SageMaker JumpStartaws.amazon.com
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
