Dear Authors,
Thank you very much for your high-quality open-source contribution. I have currently completed the training for the first 2 stages and have a few questions regarding the reproduction process:
- In the code, both Stage 1 and Stage 2 are set to train for only 1 epoch. Is this the correct configuration as intended?
- My achieved Meta Action Accuracy in Stage 1 is approximately 0.6, whereas the provided Drivema-2B model achieves around 0.7. Is this performance gap considered normal?
- My RFS score in Stage 2 is 7.28, which shows a discrepancy compared to the 7.89 reported in the paper.
I would appreciate any insights or suggestions you could provide.
Best regards
Dear Authors,
Thank you very much for your high-quality open-source contribution. I have currently completed the training for the first 2 stages and have a few questions regarding the reproduction process:
I would appreciate any insights or suggestions you could provide.
Best regards