A recent study conducted by Just Facts examined the performance of four leading AI chatbots—ChatGPT, Gemini, Grok, and Claude—on their ability to respond to politically charged questions. The study found that three of the chatbots were more accurate in answering questions designed to elicit falsehoods from the political right than from the left. ChatGPT correctly answered 94% of questions from the right and 75% from the left, while Gemini and Claude had similar results. Grok, however, performed better on left-leaning questions, scoring 84% compared to 73% on right-leaning questions.
The study also highlighted concerns regarding the validity of sources cited by these chatbots, revealing that approximately 46% of the sources were deemed legitimate. ChatGPT had the highest validity rate at 57%, while Grok had the lowest at 32%. The study cautioned that the structured nature of the questions may not reflect the chatbots' performance on broader inquiries requiring critical thinking. The findings raise questions about the reliability of AI chatbots in providing accurate information, especially in politically sensitive contexts.
Jim Agresti, president of Just Facts, emphasized the importance of verifying information provided by AI systems, likening their reliability to that of a 'B student.' The study also noted that previous research has indicated a significant percentage of AI-generated references in biomedical contexts may be fabricated, underscoring the potential risks associated with misinformation in critical areas such as public policy and health.