AMAX details three-agent on-prem AI factory workflow
AMAX Engineering
· August 21, 2026
· ✓ verified
AMAX has published a technical blog describing how it built an on-prem “AI Factory” using LiteLLM as a unified proxy for multiple local LLMs. - The post says AMAX CTO Office deployed a LiteLLM-based inference gateway on a B300 server at its Fremont HQ, routing requests to NVIDIA NIM and vLLM models with a single OpenAI-compatible endpoint and virtual API keys.
- It says the team completed the setup in a 1-week sprint, using Claude Code for server configuration and Copilot for application migration; it also describes a customer evaluation flow that could involve a $800K server purchase and a $50 spend limit with keys expiring in 48 hours.
- The article is a technical blog / operational account, not a formal product press release, and it emphasizes infrastructure, access control, token tracking, and ROI measurement for internal AI and GPU hardware use cases.