Engineering PapersSearch

SEARCH · Engineering Papers

Results for “AI Hallucinations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Limiting Hallucinations in AI LLMs

Limiting the hallucinations made by AI by creating a new way that LLMs access the data/knowledge they are given.

Hassan, Hashim [Fermilab]

Reducing AI RAG Hallucination by Optimizing Routing Techniques

Large Language Models (LLMs), such as ChatGPT, tend to “hallucinate”, meaning they confidently generate false information. Retrieval Augmented Generation (RAG) attempts to diminish hallucination by providing context to the LLM from data stores (indexes) containing relevant information. The LLM uses this context to formulate its response. RAG systems can still suffer from hallucination because of bad embeddings or ineffective routing. For example, a router will often return context from an irrelevant index, resulting in a hallucinated answer. In this study, we aim to minimize the frequency of routing hallucinations by optimizing Index Summary Routing.

97 MATHEMATICS AND COMPUTING

Intern Poster

Large Language Models (LLMs) have skyrocketed in popularity after the release of ChatGPT in late 2022. Although LLMs are powerful tools, they can be subject to hallucinations, which is when an LLM (or any AI model) produces misleading/ nonsensical information. The objective is to determine if statistical methods can be used to detect hallucinations as an LLM generates its answer token by token (essentially word by word).

97 - MATHEMATICS AND COMPUTING

REFSafE: A RAG-Enabled Framework for Predictive Risk Analysis and Automated Safety Report Generation in Mission-Critical Environments

Operational safety in mission-critical environments requires AI systems that are accurate, interpretable, and resistant to hallucination. We present an agentic Retrieval-Augmented Generation (RAG) framework, REFSafe, for grounded hazard analysis and automated safety report generation. The system integrates Large Language Models (LLMs) with structured operational data, historical incident repositories, policy documents, and external authoritative sources. Through iterative agentic reasoning, the framework retrieves, verifies, and synthesizes evidence prior to generation, enforcing citation-backed outputs with explicit source attribution (documents, links, and prior events) to ensure traceability and trust. To mitigate hallucinations and unsupported claims, all risk assessments and forecasts are constrained to retrieved evidence, with confidence signals derived from retrieval relevance and source consistency. A transparent pipeline enables subject matter experts (SMEs) to validate predictions, and provide structured feedback, forming a continuous performance calibration loop. Preliminary deployment demonstrates improved reliability in hazard detection and safety/vulnerability report generation. This work advances trustworthy, evidence-grounded AI for predictive safety intelligence in mission-critical operations.

Das, Sanjay [ORNL] (ORCID:0009000542591915)

Harnessing the Power of AI: Status and Expansion of Current Domestic Transport Security Through Flexible Embedded Hardware

As applications of Artificial Intelligence (AI) continue to expand, there are increasing opportunities to leverage applied AI methodologies with mobile transportation focused embedded systems. Current applications of AI in transportation focus on a variety of areas, including fuel efficiency, safety, security, and other broad fields of optimization or detection. To leverage these AI workflows and methodologies in the field, teams must utilize complex embedded systems capable of implementing these AI-enabled algorithms in real-time. In this paper, we will investigate how these algorithms can be integrated into existing technologies leveraging vehicle data - such as the Controller Area Network Transport Security Tracking and Reporting Unit (C-STAR). The C-STAR technology is an embedded platform with onboard computation capable of running next generation algorithms in vehicle systems AI, such as preventative maintenance, driver authentication, and transport security. As deployed in the field, the C-STAR has a limited AI functionality –this paper will directly discuss how a device like C-STAR can be utilized and the advantages of integrating these new technologies. We will open with relevant background information and transportation projects that leverage AI, focusing specifically on those around transport security such as vehicle identification, anomaly detection, and deterrence. We will then extend this into potential opportunities and scaling for AI methodologies using platforms like the C-STAR. Finally, we will speak directly to the challenges of deploying AI-powered workflows, such as computing power needs, bandwidth, hallucinations, and other regulatory considerations.

Cook, Adian [ORNL] (ORCID:0000000160825395)

AstraAI v1

AstraAI is an open-source, structure-aware AI coding agent designed for large scientific and DOE-HPC codebases such as AMReX-based applications. Unlike general-purpose coding assistants, AstraAI combines retrieval-augmented generation (RAG) with compiler-level Abstract Syntax Tree (AST) analysis to perform precise, scope-constrained code modifications. It identifies exact function spans, enforces locality of edits, and maintains cross-file invariants, enabling deterministic and build-safe transformations in complex C++/GPU environments. AstraAI is intended for developers working on large, evolving HPC frameworks where correctness, reproducibility, and structural integrity are critical. Typical use cases include modifying physics kernels, updating GPU device lambdas, and performing multi-file refactors without breaking compilation or runtime semantics. Compared to conventional LLM-based coding agents - even those with repository access - AstraAI provides structural guarantees rather than free-form text patches. It minimizes unintended diffs, prevents scope drift, preserves formatting and build stability, and reduces structural hallucinations. By integrating compiler tooling directly into the generation loop, AstraAI transforms AI-assisted coding from probabilistic text editing into deterministic, structure-preserving program transformation suitable for mission-critical scientific software.

Natarajan, Mahesh [Lawrence Berkeley National Labo

Privacy-Aware RAG-Enabled LLMs for Collaborative AI in Organizations

Recent advancements in Large Language Models (LLMs) based on Transformer architectures have significantly improved capabilities in natural language processing and generation. However, deploying LLMs for inter-organizational communication poses challenges, in ensuring privacy and facilitating effective collaboration. This paper introduces a novel decentralized inference meta-agent chatbot that leverages privacy-aware Retrieval-Augmented Generation (RAG)-enabled LLMs for collaborative AI communication across organizations. Built on Microsoft’s Autogen, the platform enables LLMs to autonomously refine responses, enhancing accuracy and relevance. It incorporates advanced hallucination mitigation techniques using Uptrain and a privacy-focused RAG framework that employs synthetic document generation to protect sensitive information. Comprehensive evaluations demonstrate the platform’s effectiveness in maintaining contextual relevance and stringent privacy standards, effectively addressing critical challenges in LLM-enhanced collaborative AI communication. This work represents a significant step toward secure and efficient inter-organizational collaboration using advanced generative AI technologies.

97 - MATHEMATICS AND COMPUTING

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

J Lemery

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

Medical Operations

AI for Interpreting Nuclear Power Plant Documents for Power Uprates

To reduce the cost and time needed for regulatory compliance, nuclear power plants (NPPs) can utilize artificial intelligence (AI) to assist in interpreting complex and voluminous documents that typically span thousands of pages. Usually, the process of interpreting a plant’s technical specifications (TSs) and associated documents is labor intensive. This study aims to understand what processes state-of-the-art large language models (LLMs) can automate and to identify the pitfalls associated with using LLMs to reduce human labor costs and time. This research uses a recent AI technology called retrieval augmented generation (RAG), which retrieves pages of information from TSs and associated documents to assist with NPP power uprates (cleared to produce more power). LLMs are integral to RAG because they create human-like responses based on the retrieved information, aiding in the interpretation and application processes. A baseline case demonstrates how LLMs can operate successfully for a power uprate application. Then five use cases show five types of potential failures: (1) RAG retrieving the incorrect information, (2) RAG misinterpreting the retrieved information, (3) RAG relying on knowledge not contained in the retrieved information, (4) RAG hallucinating, and (5) RAG refusing to answer. The results of the five use cases suggest that automating the human interpretation of TSs and associated documents with AI should be approached with caution. A subject-matter expert reviewed the AI outputs from the five use cases and concluded that an LLM can produce technical information that is needed to produce power uprate applications in certain instances.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

From Rules to Reasoning: A Survey of Large Language Model-Based Approaches to Scientific Hypothesis and Idea Generation

Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.

AI-driven discovery

chatHPC: Empowering HPC users with large language models

The ever-growing number of pre-trained large language models (LLMs) across scientific domains presents a challenge for application developers. While these models offer vast potential, fine-tuning them with custom data, aligning them for specific tasks, and evaluating their performance remain crucial steps for effective utilization. However, applying these techniques to models with tens of billions of parameters can take days or even weeks on modern workstations, making the cumulative cost of model comparison and evaluation a significant barrier to LLM-based application development. To address this challenge, we introduce an end-to-end pipeline specifically designed for building conversational and programmable AI agents on high performance computing (HPC) platforms. Our comprehensive pipeline encompasses: model pre-training, fine-tuning, web and API service deployment, along with crucial evaluations for lexical coherence, semantic accuracy, hallucination detection, and privacy considerations. Here, we demonstrate our pipeline through the development of chatHPC, a chatbot for HPC question answering and script generation. Leveraging our scalable pipeline, we achieve end-to-end LLM alignment in under an hour on the Frontier supercomputer. We propose a novel self-improved, self-instruction method for instruction set generation, investigate scaling and fine-tuning strategies, and conduct a systematic evaluation of model performance. The established practices within chatHPC will serve as a valuable guidance for future LLM-based application development on HPC platforms.

97 MATHEMATICS AND COMPUTING

Agentic framework for programmatic crystal structure generation using a fine-tuned worker–supervisor large language model

Platinum group metals (PGMs) underpin many catalytic technologies but face severe supply constraints, motivating the search for alternative materials and computational methods to accelerate discovery. While atomistic simulation tools such as Pymatgen and ASE have streamlined structure manipulation, they require detailed inputs, limiting accessibility for experimentalists and slowing early-stage exploration. Here, in this study, we present an AI-driven agentic framework that orchestrates worker–supervisor large language models (LLMs). The worker translates natural-language prompts of varying abstraction into valid crystallographic structures using a compact LLM fine-tuned with low-rank adaptation on a curated text–code–CIF dataset, emphasizing energy-efficient training. Benchmarking against the baseline CodeGen-350M-mono model shows that fine-tuning reduces hallucination rates from 100% to as low as 5% and improves structural match accuracy to up to 82% for fully specified inputs. Accuracy declines with decreasing prompt detail but remains nontrivial even when only stoichiometry and space group are provided, underscoring the LLM’s capacity for crystallographic inference. The supervisor Claude LLM evaluates the outputs and triggers iterative refinement through the worker’s built-in structure manipulation capabilities (e.g., supercell scaling, strain, vacancy, and substitution operations). We further demonstrate use cases for technologically relevant catalysts, including IrO 2 , pyrochlore Pb 2 Ir 2 O 7 , Ni 2 FeO 4 , and Ni 3 Mo, where the framework generates physically consistent structures that can be refined via geometry optimization. This work introduces a low-energy, language-driven pathway for integrating human and machine intelligence in materials design, paving the way for AI-assisted synthesis planning and high-throughput screening of complex oxides.

AI agent

Role of Uncertainty Quantification in the Explainability of Large Language Models for the Nuclear Industry

The meteoric rise of generative artificial intelligence (AI) large language models (LLMs) has created an opportunity to utilize them to increase efficiencies in a multitude of industries. While LLMs carry great potential to revolutionize the manner in which work is performed, numerous known deficiencies limit their utility, including the black box nature of the models, the stochastic nature of the response (i.e., presenting the same prompt multiple times results in different responses), and the potential for hallucination. Widespread adoption of LLMs in safety-critical industries such as nuclear will require some form of explainability to assure end users that the LLM’s response to a given query is valid. Model uncertainty is inherently linked to the concepts of trust and explainability, and can be used to identify situations in which the model is insufficiently certain about its answer. Although uncertainty is not enough in and of itself to determine the suitability of an answer—a model can be very certain of an inaccurate answer—it still provides valuable supporting information. Practical methodologies for gauging or quantifying the uncertainty in LLM outputs are presented herein, along with examples based on nuclear-specific prompts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Atlas: Navigating NASA’s Knowledge Universe with AI-Powered Natural Language Queries

NASA has a vast archive of engineering guidelines, standards, and best practices collected over decades. This encompasses a breadth of topics from rocketry and engineering standards to risk management and space-related health issues. This wealth of information, while invaluable to NASA engineers, staff, and the public, is too extensive for any individual to fully comprehend. To address this challenge, we have developed Atlas, a tool within NASA's Mission Cloud Platform that enables users to query these diverse sources effectively. Atlas allows users to ask natural language questions and receive answers grounded in factual information from source documents. The tool provides responses with direct quotations and links to original documents, ensuring transparency and accuracy. It can address a wide range of queries, from specific technical details like safe distances for rocket launches from lightning to broader topics such as crew health requirements for long-duration space missions, corrosion protection in low Earth orbit, and NASA's agreements with various entities. In developing Atlas, we encountered and overcame several technical challenges. Large Language Models often struggle with consistently providing accurate information, especially for highly specialized topics. We implemented strategies to prevent hallucinations and ensure the reliability of responses, even for complex questions on topics ranging from NASA Mission Classes to intricate rocket science concepts. Additionally, we addressed the challenges of delivering quick responses while maintaining cost-effectiveness. Our presentation will detail the innovative approaches we employed to optimize performance and efficiency, making Atlas a powerful and practical tool for accessing NASA's extensive knowledge base.

Artificial Intelligence

Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning

Despite the dramatic progress in Large Language Model (LLM) development, LLMs often provide seemingly plausible but not factual information, often referred as hallucinations. Retrieval-augmented LLMs provide a non-parametric approach to solve these issues by retrieving relevant information from external data sources and augment the training process. These models helps to trace evidence from an externally provided knowledge base allowing the model predictions to be better interpreted and verified. In this work, we critically evaluate these models in their ability to perform in scientific document reasoning tasks. To this end, we tuned multiple such model variants with science-focused instructions and evaluated them on a scientific document reasoning benchmark for the usefulness of the retrieved document passages. Our findings suggest that models justify predictions in science tasks with fabricated evidence and leveraging scientific corpus as pretraining data does not alleviate the risk of evidence fabrication.

• Artificial intelligence (AI) / machine learning