To be trustworthy, LLMs need to show their work
Large language models (LLMs) used in chemistry and drug discovery often generate plausible but incorrect or "hallucinated" information. To be trustworthy, these AI systems need to cite actual sources and show their reasoning, rather than simply producing confident-sounding answers without evidence. Researchers are working on methods to make LLMs more transparent and accountable for their outputs.
Background
- The article discusses "trustworthy AI" in chemistry and drug discovery, arguing that large language models (LLMs) must provide transparent reasoning ("show their work") rather than just outputting answers.
- Large language models (LLMs) are AI systems trained on vast text data that can generate human-like responses; they are increasingly used in scientific research to propose molecules, predict reactions, or summarize papers.
- Critics note that LLMs often "hallucinate" — confidently producing factually wrong or nonsensical outputs — which is especially dangerous in drug development, where mistakes could lead to toxic compounds or wasted resources.
- The piece calls for integrating explainability tools (like chain-of-thought reasoning or citation generation) so chemists can verify AI suggestions before trusting them.