Error benchmark for AI peer review: 100 known errors in 10 psychology papers, scored across 14 LLM and commercial reviewer configurations
Reddit r/datasets1mo4 min read
submitted by /u/cavedave [link] [comments]
submitted by /u/cavedave [link] [comments]