StepFun: Step 3.7 Flash
π§ AI Modelstepfun
196B MoE multimodal model, 11B activated, 256k context, low cost.
Step 3.7 Flash is the latest multimodal MoE model from StepFun. It combines a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token for high efficiency. The model supports a 256k token context window, enabling processing of long documents or high-resolution videos. It achieves ELO scores of 1195 in 3D understanding, 1211 in ASCII art, and 1220 in code categories on OpenRouter benchmarks. The model is available via OpenRouter with structured outputs, reasoning, and logprobs support.
π‘Highlights
- ββ196B MoE, 11B activated per token
- ββ256k context length
- ββText, image, video input
π―For
- ββAI researchers
- ββdevelopers
- ββenterprises
πLinks
- ββOpenRouter Model Page