▲ 4 points
back
1 comments
"Orca surpasses ... Vicuna-13B by more than 100% in complex zero-shot reasoning benchmarks like Big-Bench Hard (BBH) and 42% on AGIEval. ... reaches parity with ChatGPT on the BBH benchmark and shows competitive performance in ... SAT, LSAT, GRE, and GMAT, while trailing behind GPT-4."