A wonderful job!!! I am using the benchmark provided by you to test my Reward model. But I have encountered some problems with the pairwise comparison data. Specifically,I used the test code provided in your github repository[scripts/gen_batch_api_judge.py]. But I found that this code only generates Judge responses for pair data. I'm not sure which answer is better so I have no way to calculate the score of the test result. I checked the entire code repository but didn't find the file storing this information. Could you help me? Thank you very much.
A wonderful job!!! I am using the benchmark provided by you to test my Reward model. But I have encountered some problems with the pairwise comparison data. Specifically,I used the test code provided in your github repository[scripts/gen_batch_api_judge.py]. But I found that this code only generates Judge responses for pair data. I'm not sure which answer is better so I have no way to calculate the score of the test result. I checked the entire code repository but didn't find the file storing this information. Could you help me? Thank you very much.