MiniMax M2.7 is a model used by many of our customers around the world that helps augment their coding and agentic workflows using the fast inference speed of SambaNova’s SN40 to accelerate their tasks. SambaNova’s SN50 chip improves running this model significantly, which we previewed at RAISE with MiniMax M2.7 running over 800 tokens/second as validated by Artificial Analysis.
SemiAnalysis has now benchmarked SambaRack SN50 on MiniMax M2.7 as well, demonstrating two configurations of the same model with different chip parallelism strategies. In the chart below, SN50 is deployed using vLLM appearing with ~800 tokens/second at the fastest interactivity and ~400 tokens/second with throughput starting to approach the fastest throughput of B200s running the same model. This is early validation for SambaRack SN50 on MiniMax M2.7, showing fast decode performance in the part of the stack users feel the most.
Speed Users Feel, Balanced with Throughput Inference Providers Need
This chart matters because it shows the two things inference providers have to optimize at the same time. The x-axis is interactivity, or tokens per second per user, which is the speed developers feel when a coding assistant or agent is generating responses in real time. The y-axis is output token throughput per chip in the chart’s normalization, which points to how much work each accelerator can sustain as providers scale demand. Moving right means a more responsive experience; moving up means more throughput capacity by serving more requests.






