VKAE: VIDRAFT's Inference Engine Hits 23× GPU Speedup and ~10K Tokens/sec on a Single Nvidia B200
TL;DR: VIDRAFT, a Korean Pre-AGI AI startup, has developed VKAE — an inference acceleration system that achieves up to 23× GPU utilization improvement on a single Nvidia B200, delivering approximately 10,000 tokens per second on the Qwen3.5-35B-A3B model. The system exposes an OpenAI-compatible API, making it a drop-in target for existing LLM toolchains.
What it is
VKAE is VIDRAFT's proprietary inference engine designed to dramatically increase the throughput efficiency of large language model (LLM) serving on modern accelerator hardware. According to reporting from Russia's RusNews, the system was benchmarked on a single Nvidia B200 GPU and demonstrated:
Up to 23× improvement in GPU utilization compared to baseline inference configurations







