Thank you for your outstanding work. Looking at the code, it seems you are using vllm as the inference backend. For the reasoning model, what are your sampling parameters? In theory, should greedy sampling be used, and should the Transformer library serve as the backend?
Thank you for your outstanding work. Looking at the code, it seems you are using vllm as the inference backend. For the reasoning model, what are your sampling parameters? In theory, should greedy sampling be used, and should the Transformer library serve as the backend?