ParaTrace AI Text Detector vs Pangram: Re-analysis of AI-Generated Paper Detection in NeurIPS Submissions
By Mr.Xu
Published:
Summary:The ParaTrace team reanalyzed the AI-generated text detection results for the NeurIPS 2026 paper submissions and compared them with Pangram 3.3.2. The study reveals discrepancies between the two detectors, with ParaTrace flagging more AI-generated texts in 2025 submissions while Pangram remains more conservative. This disparity may stem from ParaTrace's sensitivity to AI-assisted polishing and NeurIPS's policy allowing AI editing. ParaTrace outperforms Pangram in detecting clean AI text but lags
Background and Motivation
During the review process of the NeurIPS 2026 submissions, the chairs used Pangram 3.3.2 to detect AI-generated text in papers from an AI ethics conference, which included submissions from both 2022 (before ChatGPT) and 2025. The ParaTrace team independently reanalyzed these papers using ParaTrace v7.
Key Findings
-
2022 Papers
- Both detectors found no AI-generated text in the 2022 papers.
- ParaTrace flagged 0.4% of the 4,512 text passages as AI-generated.
- With a sample of 106 papers, the upper bound of the false-positive rate at 95% confidence is approximately 3.5%.
-
2025 Papers
- At ≥90% AI-generated text, both detectors flagged 2 papers.
- At ≥50%, ParaTrace flagged 9 out of 123 papers, while Pangram flagged about 2 out of 204.
-
Analysis of Discrepancies
- AI-Assisted Writing: Pangram may not detect some AI-assisted writing.
- AI-Polished Text: ParaTrace counts AI-polished text as AI-generated, while NeurIPS allows AI editing, which may lead to false positives.
- Text Chunk Size: ParaTrace uses 512-character chunks for analysis, whereas Pangram's approach may differ.
Technical Highlights
- ParaTrace's Strengths: Excels in detecting clean AI text with a 99.6% detection rate and 98.6% on unseen 2025 model versions.
- ParaTrace's Weaknesses: Less effective in handling short texts and humanized content, with a higher false-positive rate.
- Pangram's Strengths: More stable in detecting short texts and humanized content, with a lower false-positive rate.
Industry Implications
- Complementarity of AI Text Detectors: Different detectors have unique strengths and should be combined for more accurate and reliable detection.
- Impact of NeurIPS Policy: The policy allowing AI editing may affect the accuracy of detectors, highlighting the need for clearer boundaries in AI ethics and academic integrity.
Developer Recommendations
- Use Multiple Tools: Combine multiple AI text detection tools to improve accuracy and reliability.
- Monitor False Positives: Pay special attention to false positives when dealing with short texts and humanized content, and incorporate human review.
- Continuously Optimize Models: ParaTrace, Pangram, and other tools should continue to refine their algorithms to keep up with the evolving landscape of AI-generated text.
Conclusion
ParaTrace and Pangram each have their own strengths and weaknesses in AI text detection. Developers should choose the appropriate tool based on their specific needs and scenarios and combine it with human review to ensure the accuracy of detection results.
— END —Source: Reddit r/MachineLearning (2026-10-08)
Tags: #AI Text Detection #ParaTrace #Pangram #NeurIPS #AI Ethics
Community Comments