Which models can your machine actually run? Magnitude is a 100% free, open source desktop app that: - profiles your hardware and runs sample calculations - predicts tok/s for every model before you download - recommends the best models, from fast to smart Pick your models and it handles the rest: - downloads and tunes the models for your hardware - connects agents like pi, opencode, hermes in one click - runs models on demand as your agent works Works on whatever hardware you already own: MacBook (Intel or Apple Silicon), Mac Mini, Mac Studio, DGX Spark, Strix Halo, any NVIDIA/AMD GPU, or just a CPU Download it on macOS, Windows, or Linux: Open source:
Organisation
Apple
The device maker, whose Apple Intelligence features run models on-device and route the rest through a private cloud it operates itself.
Apple was recorded in 24 items across 7 of the 8 briefings in the current window.
Its share of coverage was steady: 9 items in the first half of the window and 15 in the second, tracking the feed as a whole, which grew about 2.2×.
It appeared most often alongside NVIDIA, GitHub and DGX Spark.
- items
- 24
- briefings
- 7
- mentions
- 30
- last seen
- 2026-09-19
- Official channel
- apple.com
- Background
- en.wikipedia.org
Coverage timeline
Sat 12 Sept – Sat 19 Sept / 8 briefings
Appears alongside
NVIDIA
OrganisationThe accelerated-computing company whose GPUs and DGX systems train and serve most frontier models. Its data-centre roadmap is a constraint the rest of the field plans around, because training capacity depends on it.
25 items / 8 briefings
GitHub
OrganisationThe code hosting platform owned by Microsoft, and the home of Copilot — one of the earliest and most widely deployed coding assistants.
67 items / 8 briefings
DGX Spark
ProductNVIDIA's desktop AI computer, aimed at developers who want to run models locally rather than rent capacity.
16 items / 8 briefings
Claude
ProductAnthropic's assistant and the model family behind it. It is used both as a consumer product and, through the API, as the base for a large share of business tooling.
66 items / 8 briefings
Alibaba
OrganisationThe Chinese cloud and commerce group whose Qwen team publishes the Qwen model family. Qwen models are both served through Alibaba Cloud and released openly, which has made the family a common base for work done elsewhere.
14 items / 8 briefings
Claude Code
ProductAnthropic's agentic coding tool, which runs in the terminal and the IDE. It is the agent most directly comparable to OpenAI's Codex, and the two are usually assessed against each other.
81 items / 8 briefings
Everything recorded
19 September 2026 2 items
Breakthrough
@inco_aiQwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡ > Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon. > Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
18 September 2026 6 items
The demo runs on your GPU now. It wanted ~25GB of VRAM on launch day. 💻This week: CPU offload, a ~7.5GB encoder cut, an MLX branch for apple silicon.
@TencentHunyuan🚀 AuK is officially here. Nano banana🍌 for audio > An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback. > 🤗 Paper & upvote: ⭐ GitHub & star:
This is going to change everything 🚀 DGX Sparks and Mac Studios working together for optimal performance. We are so early.
@volatilemarktsIT WORKS! > For a year, anyone with DGX Sparks and Mac Studios has lived with the same problem: the NVIDIA boxes are fast at reading, the Apple boxes are fast at answering, and they can't share a thought. Two islands. A 10 GbE cable between them.
Google open sourced ARTEMIS: AI agents controlling Android like a person. Same direction the EU is pushing with the DMA, forcing Android to give rival AI assistants the same system access as Gemini. Apple is fighting that battle over Siri. Either way, the OS stops being a moat.
录完视频,剩下的多机位导播、加字幕和跨平台发布,现在可以全部丢给 Claude Code 处理了。 VibeTube 是一个开源的 macOS 录屏工具。它的核心逻辑非常纯粹:“你只管录,AI 负责剪辑和发布”。工具会将同步好的屏幕和摄像头素材,直接交给本地运行的 Claude Code 或 Codex 进行自动化后期。 • AI 自动导播:无需手动打轴,AI 代理会根据你的讲解内容,自动完成镜头选择与机位切换(全尺寸人像 / 屏幕录制 / 画中画),并配上字幕与音效。 • 内置影音增强:集成 NVIDIA Studio Voice NIM (48k-hq) 消除房间混响与底噪;利用 MatAnyone2 (Apple Silicon) 直接在本地完成背景替换。 • 零干预发布:AI 会读取最终的成品字幕,生成 5 种不同视角的备选标题以及带真实时间戳的 YouTube 章节描述,最后直接推送至 YouTube、TikTok 和 Reels,全程无需打开浏览器。 适用限制与门槛: 目前仅支持 macOS 环境。需本地安装 Node.js 22+ 与 ffmpeg,并自备对应的 CLI 工具与 API 密钥(Claude/Codex、NVIDIA Studio Voice 及 Upload-Post)。
Open Research is really unbeatable!
@gajeshTogether, the MLX(.)fast community has made Qwen 3.8 Flash nearly 2x faster on Apple Silicon! > We're ready to bring it to @DarkbloomAI: an open network of local Mac machines providing inference to the world. One thing remains: the community flagged that its license requires a separate agreement for commercial model serving, so we're holding the launch until that's in place. > We believe this is a great opportunity for the local community: one where we make Qwen models faster and more accessible, and the people running them share in the value they create. > .@Alibaba_Qwen @QwenDevs, we'd love to work together on this. > If anyone else knows someone we can talk to, we'd love to have that conversation. Let's make Qwen 3.8 Flash on Darkbloom a reality!
This is insane! Codex Remote Control for Apple Watch:
17 September 2026 3 items
Local model ship on your iPhone and Mac natively
@LocallyAIAppTry the new Apple Foundation Models, available in the app on iOS 27. > Updated with better answers, improved instruction-following, and now image understanding.
OpenClip 是一个 macOS 上的开源小工具,在任何应用里选中一段文字,旁边就浮出一条操作栏,用过 PopClip 的一看就懂。 复制、搜索、大小写转换这些直接点,选中的是算式就地出结果,选中一段英文能就地总结、翻译或改写。 AI 那部分可以走 Apple Intelligence、本地的 Ollama 模型,也能接 OpenAI 或 Claude,结果卡上有替换和复制两个按钮。 GitHub: 扩展是它的重头,一个描述文件加一个脚本就是一个扩展,JavaScript、AppleScript、命令行脚本、网址模板都行,不用编译。内置了扩展商店,一键装。 也能在设置里直接加一个搜索网址或者一段脚本当动作,不用写描述文件。还能按应用定规则,比如某个动作只在终端里出现。 第一次启动有 4 步引导,授一个辅助功能权限、装几个基础扩展,最后给个练手区试一遍。 Homebrew 一条命令装好,要 macOS 14 以上。每天在 Mac 上复制来复制去切窗口的,装上能省不少功夫。
Incredible work! As soon as I can get my hands on a DGX Spark, I'm combining this with my work on @OmarchyMac and omarchy-mlx to bring this to Linux on Apple hardware. Who do I know that has good connections at NVIDIA to make this happen?
@ashxhartMCDMA 0.1.18 is out ✅ > Larger Registered Buffers Teardown fixes A CLI tool Bug fixes >
16 September 2026 4 items
Ian Failes from befores & afters chats with Nikola Todorovic, who co-founded Wonder Dynamics, an Autodesk Company, about the new 3D Editor + Canvas in Flow Studio. Spotify: Apple Podcasts:
Apple's MobileCLIP2 matches models 2.3x its size - fast image-text understanding built for the edge, not the data center. On-device multimodal AI just got a serious upgrade. Github:
视频生成加速框架 FastVideo 最新释出 FastH3 8-Step V2,并全面打通 Mac 本地 MLX 推理链。 它不是单纯的模型搬运库,而是一套覆盖分布式微调与端到端优化的完整工作流(目前 4.4k Stars)。对于想要在本地完成高品质视频生成的开发者,这次更新直接命中了算力和硬件门槛的痛点。 核心工程进展: • 算力开销大幅压缩:新发布的 FastH3 8-Step V2 基于 MiniMax-H3 进行 DMD2 步进蒸馏,引入高达 80% 的视频稀疏注意力(Video Sparse Attention),极大降低了推理成本。 • Apple Silicon 原生支持:告别云端依赖。借助 MLX 框架与 FastMetal-QAD,Mac 用户现在可以原生运行从 1.3B 到 14B 参数的视频生成模型。 • 多端适配与实时编辑:除了主流 NVIDIA 显卡,现已支持 DGX Spark 环境(注:ARM64 架构目前暂无预编译 wheel,需从源码编译 CUDA kernel)。其内置的 Dreamverse 模块可实现本地视频流的实时“Vibe Directing”控制。 如果你习惯在 macOS 桌面上(例如 32GB 统一内存的 Mac 环境)进行模型部署与自动化测试,FastVideo 提供的清晰 CLI 与 Python API 能帮你快速跑通从代码到视频的最后一公里。 项目文档与源码:
Deepseek 4.1 flash uncensored
@OrcaRouter🐳 Run DeepSeek V4.1 Flash locally on your Mac — uncensored for security research. > We just released Orca’s official MLX weights for Apple Silicon. > This build is designed for AI security research, red teaming, alignment research, and agent-security testing — where refusal behavior itself can get in the way of measuring the model. > 4-bit — recommended → 458.7 GB → 0.9954 routed-expert fidelity → 512 GB Mac > 3-bit → 364.3 GB / 512 GB Mac 2-bit → 212.2 GB / 256 GB Mac > On our refusal evals, the uncensored build reduced harmful-prompt refusal by 87–96% across JBB, AdvBench, MaliciousInstruct, HarmBench, ForbiddenQuestions, StrongREJECT, and SimpleSafetyTests. > Built for security researchers who need to study what happens when the guardrails come off. > The downloadable weights are uncensored. Our hosted API remains guardrailed.🐳 > API:
15 September 2026 6 items
Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.📜 Apache 2.0. 🤖 ⚡ Reaches 23.9–25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark. 🏆 Retains most of the FP16 base model’s quality, with an average gap of just 2.8 points across five benchmarks. MMLU-Pro rises from 65.8 to 70.1. 🧠 SSD expert offload keeps most weights outside active memory, while a prerouter predicts which experts each layer will need next. 🛠 Recover-LoRA offsets quality loss from 4-bit quantization and expert offloading. The current Preview runs locally through MLX on Apple Silicon.
A Metal capability layer for macOS VMs just unlocked 11-16x faster LLM inference on Apple Silicon - no hardware changes, just a process-scoped shim hitting the paravirtual GPU. Gemma 4 12B goes from 3.41 to 49.67 tokens per second on generation. Local AI in VMs just became a serious option.
@trycua1/ Today, as part of our broader research into Apple Silicon virtualization, we're releasing a process-scoped Metal capability layer for macOS VMs. On one M1 Ultra, prompt / generation: > TinyLlama: 11.08× / 16.36× Gemma 4 12B: 7.20× / 14.54× Muse Glimmer 30B: 7.55× / 8.87×
Wow... this is very cool.
@tonysimons_I just found a repo that lets AI agents control a VIRTUAL IPHONE. > Not the iOS Simulator. > Actual virtualized iOS on Apple Silicon. > Screenshots. Touch. Swipes. Typing. Apps. > There’s even an MCP server for coding agents. > +633 stars TODAY. >
DeepSeek Harness 官方桌面端已经基本做完。官方仓库主分支新增完整的 apps/desktop,用 Electron 把现有 Harness Web UI 封装成桌面 App。应用自带 Node.js 和 pnpm,用户不需要再手动启动 Web 服务,桌面端也不会对外监听端口。 官方已经准备好 macOS Apple Silicon、Intel 和 Windows x64 三套安装包构建流程,还包括 macOS 签名与公证、Windows EV 签名、自动更新和安装失败恢复。生产更新服务器已经指向 DeepSeek 自己的 暂时不在官方发布目标里。 最近一轮桌面端提交主要在处理 macOS 公证提速、启动恢复、打包校验等发布前问题。 DeepSeek 目前还没有公布具体发布日期,GitHub Release 里也还没有桌面安装包。但从仓库状态来看,桌面端已经进入发布前收尾阶段。
Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite. This architecture is a great demonstration of how to connect Slack to a project without exposing the local server! 🔗 Blog: 🔗 Repo:
yeah.. maybe make an hermes as a hidden model-provider support @Teknium @NousResearch . We need to get away from cloud sources and able to use open local models vis Hermes agent CLI integration to apple Siri.
@marcelpociotI got Claude answering inside Siri on macOS 27 🚀 > Apple's hidden model-provider support + my existing Claude Code account. > Open-source proof of concept. Requires disabling SIP/AMFI. > Thanks @itspdfu for the discovery! > Check it out:
14 September 2026 2 items
平时在 Mac 上干活,偶尔要回家里那台 Windows 上取个文件、跑个只有 Windows 才有的软件,商业远程软件不是限速就是要买会员。 YourDesk 是一款 macOS 和 Windows 互连的远程桌面工具,输入对方 ID 和密码就连上,源码在 GitHub 公开,个人和公司内部用都免费。 编解码走硬件加速,Mac 上 H.264 和 HEVC,Windows 上 H.264 和 AV1,两端按各自能力自动选,工具栏能看到当前用的是哪种。 跨电脑复制粘贴是我觉得最实用的一点,文字、图片、文件、整个文件夹都能复制过去直接粘贴,一批最大 2GB。 GitHub: 多显示器在工具栏中间直接切,最多 4 个屏。远端分辨率高的话可以开画面增强,先压低传输再本地放大,4K 桌面效果更明显。 新版加了 MCP,AI Agent 能自己连上远程电脑,操作桌面或命令行,干完活断开。这个功能默认关着,开了也只放行本机来源。 除了桌面模式还有远程命令行模式,Linux 机器也能当命令行被控端。Mac 端是签过名过了苹果公证的 DMG,Apple Silicon 专用。
Open weights. Shared progress. MiniMax H3 is moving fast. We built MiniMax H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent highlights: • FastH3 — FastVideo, Nuva Lab and NVIDIA: 4-step distillation, now running on DGX Spark and Apple Silicon. • Sol-H3 — NVIDIA’s SANA team: now on DGX Spark with a two-stage H3 + LTX-2.5 pipeline. On 8×B300, the team reports 15 seconds of 768p video + audio in 6.6 seconds of warm inference.* • VDN — Haocheng Xi and the OpenVDN team: rethinking attention for faster H3 inference, with weights, training and inference code released. • PDD — NVIDIA’s distillation method, brought to H3 by Alibaba PAI as 8-step Acc-LoRAs, now supported in ComfyUI. • LightX2V — 4- and 8-step Turbo LoRAs, with workflows for text, image and reference-conditioned video + audio. Behind every release are people training, optimizing, quantizing, testing and sh
12 September 2026 1 item
MiniMax H3 is now on my preferred image and video generator on Apple Silicon! I need to try it!
@drawthingsapp🚨 MiniMax H3 is RIGHT HERE in Draw Things!Go try it NOW! > 🔄 Draw Things v26.0910.1 was released in the iOS / macOS AppStore 2 hours ago. This version brings: > 🔹 Support MiniMax H3 series models, including LoRAs and TeaCache; 🔹 Support importing Krea 2 series models. 🔹 Fix some performance issues on M4 Apple Neural Engine. > gRPServerCLI and draw-things-cli both are updated to 26.0910.1 with above related updates.