Free tool
Local Model Advisor
Thinking about running AI on your own hardware? Tell us your team size, workload, and budget — get a concrete recommendation: which box to buy, which open model to run on it, and what software stack to serve it with. Or start from a machine you're eyeing and see everything it can run.
Local Model Advisor — recommendation
On-prem AI hardware and open-weight model sizing · agentsofwork.ai
Your needs
Our pick
$5,299NVIDIA Nemotron 3 Super on Apple Silicon Mac Studio
NVIDIA · USArtificial Analysis Intelligence Index 36 — ahead of gpt-oss-120b at 33.3, and the strongest US-developed model that fits a single 96GB card at Int4. LatentMoE with multi-token prediction (built-in speculative decoding); served by vLLM, TensorRT-LLM, SGLang or llama.cpp.
- Speed
- ~53 tok/s (estimated)
- Precision
- 4-bit (FP4/Int4)
- Memory needed
- ~74 GB of 96 GB
- Vs. cloud
- cloud stays cheaper
This box also runs: gpt-oss-120b, Qwen 3.6 27B, NVIDIA Nemotron 3 Nano (30B-A3B), Gemma 4 31B, Devstral 2 (24B), Gemma 4 26B-A4B (MoE), +7 more — swap models any time without new hardware.
- Apple's 2026 memory cuts hurt: M4 Max caps at 64GB, M3 Ultra at 96GB; larger configs discontinued or repriced.
- Listings still disagree on what's orderable — verify configs at purchase.
- MLX/llama.cpp ecosystem, not CUDA — no vLLM tensor parallelism.
Runner-up
$5,299gpt-oss-120b on Apple Silicon Mac Studio
OpenAI · USArtificial Analysis Intelligence Index 33.3. Fits a single 96GB RTX PRO 6000 at FP4 (~193 tok/s) — the fastest-serving pick in its class.
- Speed
- ~106 tok/s (estimated)
- Precision
- 4-bit (FP4/Int4)
- Memory needed
- ~68 GB of 96 GB
- Vs. cloud
- cloud stays cheaper
This box also runs: NVIDIA Nemotron 3 Super, Qwen 3.6 27B, NVIDIA Nemotron 3 Nano (30B-A3B), Gemma 4 31B, Devstral 2 (24B), Gemma 4 26B-A4B (MoE), +7 more — swap models any time without new hardware.
- Apple's 2026 memory cuts hurt: M4 Max caps at 64GB, M3 Ultra at 96GB; larger configs discontinued or repriced.
- Listings still disagree on what's orderable — verify configs at purchase.
- MLX/llama.cpp ecosystem, not CUDA — no vLLM tensor parallelism.
NVIDIA Nemotron 3 Super on Apple Silicon Mac Studio — best quality that fits your budget at 2 concurrent requests.
Estimates are heuristic (bandwidth-bound decode, GQA-era KV cache) — verify the exact model card and current pricing before you buy. Hardware pricing as of July 2026; model catalog as of August 2026. Want the cloud-cost side of the picture? Try the AI Cost & ROI Calculator.Prepared with the free Local Model Advisor at agentsofwork.ai/app/ai-hardware-advisor
Know what you'd buy — now what?
Three ways to turn that hardware plan into working AI for your business.
Self-guided
Talk to Allie
Free to start
$99 to unlock the full written report
A live 15–20 minute working session with our AI analyst. She already has your answers — she'll dig into your biggest pain point with you, by voice or text.
Discovery Conversation
15–20 minutes with Allie in your browser. Voice or text. Allie pulls problems, not pitch — walking through your day, the tasks you dread, where work piles up, and the automations that have failed. Industry-specific probes load from the vertical you select.
AI Analysis
Your full transcript is fed to multiple AI agents. They identify 5–7 opportunities, biased toward tools from your vertical's universe. No human review at this tier — analysis is fully automated.
Your Report
An auto-generated PDF with executive summary, priority matrix, tool recommendations, a 4-day quick-start plan, and financial impact. Delivered by email and stored in your dashboard.
Self-Service Next Steps
View your report any time in the dashboard. Upgrade to the $999 strategist assessment if you want guided, human-validated recommendations, purchase AOW support if you would like us to discuss an implementation strategy or do the work.
Human-led
Strategist Assessment
$999
Report delivered within 48 hours
A real strategist runs the whole process with you end-to-end — a live discovery call, human-reviewed analysis, a custom report, and a call to walk you through it.
Discovery Call
45-minute session with an AOW strategist. Recording and transcript captured. Pull problems, not pitch — walking through your day, tasks you dread, where work piles up, and failed automations. The strategist drives the conversation and adapts in real time.
AI Analysis + Human Review
Multiple AI agents process the full transcript and identify 5–7 opportunities. Your strategist then reviews and curates the output, adds picks from their own toolkit, and refines the recommendations before report assembly.
Custom Report
Built by AOW agents, then reviewed and customized by your strategist. Executive summary, priority matrix, tool recommendations, 4-day quick-start plan, financial impact and much more. Delivered within 48 hours of the discovery call.
Review Call
30-minute session — screen-share the report, walk each recommendation, and answer three closing questions: What's most urgent? DIY or want help? What's your timeline?
Talk to AOW
Free 15-minute consult
Free
No card, no obligation
Not sure which path fits? Tell us what you're wrestling with and we'll set up a short call — no pitch, just a straight answer on whether we can help.