Modular

Modular

5|0 reviews|0 favorites
Visit website
Introduction
Unified AI inference stack from custom GPU kernels to production cloud serving on NVIDIA, AMD, and other accelerators.
Listed on
September 15, 2026

What is Modular?

Modular is a unified AI inference platform that spans from GPU kernels to cloud API endpoints. It includes MAX, a hardware-agnostic serving framework with an OpenAI-compatible API, and Mojo, a systems language for writing high-performance GPU kernels. The platform runs the same model and codebase across NVIDIA, AMD, Trainium, TPU, Qualcomm, Intel, ARM, and Apple silicon, and can be deployed in Modular's hosted cloud or in your own VPC.

How to use Modular?

  • Sign up or request a demo on modular.com to get started. Access frontier and open models through Shared Endpoints via an API, or deploy dedicated endpoints for mission-critical reliability. Run the MAX framework and Mojo self-hosted, or use Modular Cloud for fully managed, pay-by-usage serving. Use the Model Library and docs to port custom models and write custom GPU kernels.

Core features of Modular

  • Unified inference stack from GPU kernel to API endpoint
  • MAX serving framework with OpenAI-compatible API
  • Mojo systems language for high-performance GPU kernels
  • Hardware portability across NVIDIA, AMD, Trainium, TPU, Qualcomm, Intel, ARM, and Apple silicon
  • Shared and Dedicated Endpoints for frontier and custom models
  • Deployment in Modular Cloud or your own VPC
  • 1000+ supported models including DeepSeek and Kimi
  • Agentic Skills and AI coding skills for model porting

User reviews

5.0

/ 5

0 reviews

5 stars
0
4 stars
0
3 stars
0
2 stars
0
1 stars
0

No reviews yet. Be the first to write one.

Similar products

Modular Embed

Add a website badge to show your product on Hootool—help others find the right AI tool for their problem. Easy to place on your homepage or footer.

Featured on

Hootool.ai