I recently came across StepFun’s open-source Step-DeepResearch, and I think it’s well worth watching.
As the chart shows, its score on the ResearchRubrics benchmark is already close to commercial systems like OpenAI DeepResearch and Gemini DeepResearch, while its inference cost is under a tenth of theirs. That makes it one of the few deep-research models right now with near top-tier performance at a very low price. During mid-training, as training tokens increased, the model’s performance on tasks like SimpleQA, TriviaQA and FRAMES kept improving steadily. The biggest gains were on FRAMES, which requires structured reasoning, and that’s quite different from models that only strengthen “search ability.” Put simply, if Gemini or OpenAI DeepResearch are the flagship phones, Step-DeepResearch is more like getting the whole deep-research pipeline right at a mid-sized 32B, and its cost makes it a great fit for running research tasks at scale and over long periods. I’m looking forward to seeing how it gets applied.