vLLM x Novita AI: Chord W4A16 INT4 MoE Kernel for Kimi K2.x
Novita AI's Chord is a W4A16 INT4 MoE CUDA kernel for Kimi K2.x, with Humming-compatible indexed and grouped SM90 paths for vLLM serving.
Novita AI's Chord is a W4A16 INT4 MoE CUDA kernel for Kimi K2.x, with Humming-compatible indexed and grouped SM90 paths for vLLM serving.
Deploy Hermes Agent on Novita Agent Sandbox using the Hermes template. Get a persistent, self-improving AI agent running in minutes with no server setup.
Deploy OpenClaw as a persistent 24/7 AI agent on Novita Sandbox with one CLI command. No runtime limits, full model control, and multi-channel support.
StarSling uses Novita Agent Sandbox to run CI optimization agents in isolated microVMs, test changes reproducibly, and scale AI workflows across customer pipelines.
Set up DeepSeek Harness with the Novita AI OpenAI-compatible API. Install the Web UI, add a custom provider, and choose models for coding agents and research.
DeepSeek V4 Flash Vision Exp on Novita AI is an experimental vision-enabled DeepSeek V4 Flash API with 1M context, image input, and $0.44/$1.32 pricing.
Compare four Ling-3.0 models on Novita AI by fit, active parameters, context, modalities, and current pricing to choose the right serverless variant.
Send image and video-derived inputs to Ling 3.0 Flash VL through Novita AI's OpenAI-compatible chat completions API.
Explore Ling-3.0-Flash-VL on Novita AI: native image and video input, 262K context, 32K output, function calling, and free access through September 23, 2026.
See Ling 3.0 Flash Sante's serverless API specs, 256K context, function calling, and time-limited free input/output pricing on Novita AI.
Send your first Ling 3.0 Flash Sante chat-completions request with the correct Novita model ID, endpoint, Python/cURL examples, and limits.
Compare four LLM API services and when each fits: direct provider APIs, unified APIs, gateways, or self-hosted open-model endpoints.
Explore GLM-5.3 Flash on Novita AI: model ID, 1M context, 128K output, multimodal input, pricing, and where it fits for coding agents.
Explore Ling-3.0-flash-Fin on Novita AI: a finance-enhanced 124B MoE with 256K context, 32K max output, and $0 per-million-token pricing.