U.S. open-source models are quickly gaining ground. @Nvidia's newest Nemotron Ultra is fast growing on Ollama and unlocking complex, longer running tasks for developers
Token usage beginning to tip towards open models. Ollama’s @jmorgan and Peter Fenton predict that a supermajority of tokens in the future will be from open models.
Watch Ollama's @jmorgan and @peterfenton on @CNBC with @dee_bosa at 12pm PT / 3pm ET to discuss "The Post-Frontier Era" and "Open-Source AI's Breakout Moment"
Big day for Ollama! When we started, open models and the open source AI ecosystem were in their early days with few believers. Our belief in open source has never wavered. (1/2)
GLM-5.2 on Ollama's cloud just got more capacity in US & Europe! Ollama's cloud for GLM 5.2 consistently delivers between 80 to 120 output tokens per second, even during peak hours, compared to 30 to 40 tok/s on other providers. Use GLM-5.2 with Claude Code: ollama launch claude (1/2)
Congratulations to our friends at @togethercompute. Exciting moment for open models!!
Gemma 4 is now nearly 90% faster on Apple Silicon with Ollama using MLX! The speedup comes from improved multi-token prediction (MTP), now on by default for Gemma 4, with more models to come. (1/2)
Run Ornith with Ollama: ollama run ornith For coding, use it with Claude or Pi: ollama launch claude --model ornith ollama launch pi --model ornith For the more capable 35B model, use: ollama launch claude --model ornith:35b
You can use GLM-5.2 and Kimi-K2.7-Code in Codex with Ollama! ollama launch codex ollama launch codex-app
You can use GLM-5.2 and Kimi-K2.7-Coding in Codex with Ollama! ollama launch codex ollama launch codex-app
🤯 GLM-5.2 is here — built for long-horizon coding and agentic tasks, now with a solid 1M-token context. The strongest open-source coding model yet! (1/3)
Ollama now supports @cline CLI with the ability to run parallel tasks via the Kanban feature. Cline is a coding agent for your editor or terminal. It reads your repo, edits files, runs commands, and shows diffs for review. Get started: ollama launch cline
.@Kimi_Moonshot's kimi-k2.7-code is now available on Ollama's cloud! On Ollama's cloud, this model is hosted in the US on the latest NVIDIA B300 datacenter GPUs. Your data stays private and is never trained on. ❤️❤️❤️ Try the model: (1/2)
Connect to your messaging apps Connect an agent to Telegram, Discord, Slack, WhatsApp, Signal, Email, and more. One agent, one memory, every surface.
Self-learning skills Hermes generates Python skills from natural-language descriptions and improves them as you use them. Start with the 70+ skills it ships with and grow your own library around your real workflows.
Multi-agent workflows Spawn parallel subagents with isolated contexts from the desktop. Research several topics at once, run batches, and review aggregated results without cluttering your main context window.
Use Ollama with Hermes Desktop by @NousResearch. Hermes Desktop brings the same agent (its multi-agent engine, self-improving skills, and messaging integrations) into a desktop app on macOS, Windows, and Linux. Run it on Ollama using local or cloud with one command: (1/2)
Gemma 4 Quantization-Aware Training (QAT) weights are now available on Ollama! They reduce memory requirements while maintaining model quality. E2B: ollama run gemma4:e2b-it-qat E4B: ollama run gemma4:e4b-it-qat 12B: ollama run gemma4:12b-it-qat (1/2)
ollama run gemma4:12b Gemma 4 12B is updated on Ollama, and available across all platforms! Try it on: Claude Code ollama launch claude --model gemma4:12b Hermes Agent ollama launch hermes --model gemma4:12b OpenClaw ollama launch openclaw --model gemma4:12b (1/2)