Engineering PapersSearch

SEARCH · Engineering Papers

Results for “prompt engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Cyber-informed Engineering Microgrid Analysis Tool

The Cyber-Informed Engineering Microgrid Analysis Tool (CIEMAT) leverages the Department of Energy’s Cyber-Informed Engineering to prompt engineering designers and operators through an analysis of the critical functions to be supported by a microgrid installations, the criticality of those functions, the impacts of denial, disruption or misuse of those functions on the microgrid and dependent functions, and the mitigations which could best prevent impacts to those functions resulting from cyber attack. Through use of this tool, microgrid designers and operators can quickly identify appropriate engineering mitigations to limit impacts from cyber attack and functions where engineering and operational staff can prioritize and guide the application of cybersecurity protections to best support the resiliency of the system.

Wright, VirginiaL [Idaho National Laboratory (INL)

Cyber-informed Engineering Battery Analysis Tool

The Cyber-Informed Engineering Battery Analysis Tool (CIEBAT) leverages the Department of Energy’s Cyber-Informed Engineering to prompt engineering designers and operators through an analysis of the critical functions to be supported by a BESS installation, the criticality of those functions, the impacts of denial, disruption or misuse of those functions within the BESS system, and the mitigations which could best prevent impacts to those functions resulting from cyber attack. Through use of this tool, BESS designers and operators can quickly identify appropriate engineering mitigation opportunities to limit impacts from cyber attack and functions where engineering and operational staff can prioritize and guide the application of cybersecurity protections to best support the resiliency of the system.

Lampe, BenjaminR [Idaho National Laboratory (INL),

Automating and Evaluating Large Language Models for Accurate Text Summarization Under Zero-Shot Conditions

Automated text summarization (ATS) is crucial for collecting specialized, domain-specific information. Zero-shot learning (ZSL) allows large language models (LLMs) to respond to prompts on information not included in their training, playing a vital role in this process. This study evaluates LLMs' effectiveness in generating accurate summaries under ZSL conditions and explores using retrieval augmented generation (RAG) and prompt engineering to enhance factual accuracy and understanding. We combined LLMs with summarization modeling, prompt engineering, and RAG, evaluating the summaries using the METEOR metric and keyword frequencies through word clouds. Results indicate that LLMs are generally well-suited for ATS tasks, demonstrating an ability to handle specialized information under ZSL conditions with RAG. However, web scraping limitations hinder a single generalized retrieval mechanism. While LLMs show promise for ATS under ZSL conditions with RAG, challenges like goal misgeneralization and web scraping limitations need addressing. Future research should focus on solutions to these issues.

Priebe Mendes Rocha, Maria Eduarda [ORNL]

Crew Launch Vehicle (CLV) Upper Stage Configuration Selection Process

The Crew Launch Vehicle (CLV), a key component of NASA's blueprint for the next generation of spacecraft to take humans back to the moon, is being designed and built by engineers at NASA s Marshall Space Flight Center (MSFC). The vehicle s design is based on the results of NASA's 2005 Exploration Systems Architecture Study (ESAS), which called for development of a crew-launch system to reduce the gap between Shuttle retirement and Crew Exploration Vehicle (CEV) Initial Operating Capability, identification of key technologies required to enable and significantly enhance these reference exploration systems, and a reprioritization of near- and far-term technology investments. The Upper Stage Element (USE) of the CLV is a clean-sheet approach that is being designed and developed in-house, with element management at MSFC. The USE concept is a self-supporting cylindrical structure, approximately 115' long and 216" in diameter, consisting of the following subsystems: Primary Structures (LOX Tank, LH2 Tank, Intertank, Thrust Structure, Spacecraft Payload Adaptor, Interstage, Forward and Aft Skirts), Secondary Structures (Systems Tunnel), Avionics and Software, Main Propulsion System, Reaction Control System, Thrust Vector Control, Auxiliary Power Unit, and Hydraulic Systems. The ESAS originally recommended a CEV to be launched atop a four-segment Space Shuttle Main Engine (SSME) CLV, utilizing an RS-25 engine-powered upper stage. However, Agency decisions to utilize fewer CLV development steps to lunar missions, reduce the overall risk for the lunar program, and provide a more balanced engine production rate requirement prompted engineers to switch to a five-segment design with a single Saturn-derived J-2X engine. This approach provides for single upper stage engine development for the CLV and an Earth Departure Stage, single Reusable Solid Rocket Booster (RSRB) development for the CLV and a Cargo Launch Vehicle, and single core SSME development. While the RSRB design has changed since the CLV Project's inception, the USE design has remained essentially a clean-sheet approach. Although a clean-sheet upper stage design inherently carries more risk than a modified design, it does offer many advantages: a design for increased reliability; built-in extensibility to allow for commonality/growth without major redesign; and incorporation of state-of-the-art materials, hardware, and design, fabrication, and test techniques and processes to facilitate a potentially better, more reliable system. Because consideration was given in the ESAS to both clean-sheet and modified USE designs, this paper will highlight the advantages and disadvantages of both approaches and provide a detailed discussion of trades/selections made that led to the final upper stage configuration.

Davis, Daniel J.

Large language models for transportation research: Methodologies, state of the art, and future opportunities

The rapid rise of large language models (LLMs) is transforming transportation research, with significant advancements emerging between 2023 and 2025, a period marked by the inception and swift growth of adopting and adapting LLMs for various transportation applications. Despite these significant advancements, however, a systematic review and synthesis of the existing literature remains lacking. This paper aims to fill this gap by providing a comprehensive review of the methodologies and applications of LLMs in transportation. We explore key applications, including autonomous driving, travel behavior prediction, and general transportation-related queries, alongside LLM methodologies such as zero- or few-shot learning, prompt engineering, and fine-tuning. From the review, critical research gaps are identified. From the methodological perspective, many of the research limitations can be addressed by integrating LLMs with existing tools and refining LLM architectures. From the application perspective, research opportunities for LLMs to address various transportation challenges are also explored. By synthesizing these findings, this review not only presents the state-of-the-art LLM adoption and adaptation in transportation, but also proposes future research directions as well as insights and recommendations for policymakers and practitioners, paving the way for greater LLM-driven research innovations in transportation in the future.

42 ENGINEERING

Use of magnetic compression to support turbine engine rotors

Ever since the advent of gas turbine engines, their rotating disks have been designed with sufficient size and weight to withstand the centrifugal forces generated when the engine is operating. Unfortunately, this requirement has always been a life and performance limiting feature of gas turbine engines and, as manufacturers strive to meet operator demands for more performance without increasing weight, the need for innovative technology has become more important. This has prompted engineers to consider a fundamental and radical breakaway from the traditional design of turbine and compressor disks which have been in use since the first jet engine was flown 50 years ago. Magnetic compression aims to counteract, by direct opposition rather than restraint, the centrifugal forces generated within the engine. A magnetic coupling is created between a rotating disk and a stationary superconducting coil to create a massive inwardly-directed magnetic force. With the centrifugal forces opposed by an equal and opposite magnetic force, the large heavy disks could be dispensed with and replaced with a torque tube to hold the blades. The proof of this concept has been demonstrated and the thermal management of such a system studied in detail; this aspect, especially in the hot end of a gas turbine engine, remains a stiff but not impossible challenge. The potential payoffs in both military and commercial aviation and in the power generation industry are sufficient to warrant further serious studies for its application and optimization.

Pomfret, Chris J.

Research Capabilities for Oil-Free Turbomachinery Expanded by New Rotordynamic Simulator Facility

A new test rig has been developed for simulating high-speed turbomachinery shafting using Oil-Free foil air bearing technology. Foil air journal bearings are self-acting hydrodynamic bearings with a flexible inner sleeve surface using air as the lubricant. These bearings have been used in turbomachinery, primarily air cycle machines, for the past four decades to eliminate the need for oil lubrication. More recently, interest has been growing in applying foil bearings to aircraft gas turbine engines. They offer potential improvements in efficiency and power density, decreased maintenance costs, and other secondary benefits. The goal of applying foil air bearings to aircraft gas turbine engines prompted the fabrication of this test rig. The facility enables bearing designers to test potential bearing designs with shafts that simulate the rotating components of a target engine without the high cost of building actual flight hardware. The data collected from this rig can be used to make changes to the shaft and bearings in subsequent design iterations. The rest of this article describes the new test rig and demonstrates some of its capabilities with an initial simulated shaft system. The test rig has two support structures, each housing a foil air journal bearing. The structures are designed to accept any size foil journal bearing smaller than 63 mm (2.5 in.) in diameter. The bearing support structures are mounted to a 91- by 152-cm (3- by 5-ft) table and can be separated by as much as 122 cm (4 ft) and as little as 20 cm (8 in.) to accommodate a wide range of shaft sizes. In the initial configuration, a 9.5-cm (3.75-in.) impulse air turbine drives the test shaft. The impulse turbine, as well as virtually any number of "dummy" compressor and turbine disks, can be mounted on the shaft inboard or outboard of the bearings. This flexibility allows researchers to simulate various engine shaft configurations. The bearing support structures include a unique bearing mounting fixture that rotates to accommodate a laserbased alignment system. This can measure the misalignment of the bearing centers in each of 2 translational degrees of freedom and 2 rotational degrees of freedom. In the initial configuration, with roughly a 30.5-cm- (12-in.-) long shaft, two simulated aerocomponent disks, and two 50.8-cm (2-in.) foil journal bearings, the rig can operate at 65,000 rpm at room temperature. The test facility can measure shaft displacements in both the vertical and horizontal directions at each bearing location. Horizontal and vertical structural vibrations are monitored using accelerometers mounted on the bearing support structures. This information is used to determine system rotordynamic response, including critical speeds, mode shapes, orbit size and shape, and potentially the onset of instabilities. Bearing torque can be monitored as well to predict the power loss in the foil bearings. All of this information is fed back and forth between NASA and the foil bearing designers in an iterative fashion to converge on a final bearing and shaft design for a given engine application. In addition to its application development capabilities, the test rig offers several unique capabilities for basic bearing research. Using the laser alignment system mentioned earlier, the facility will be used to map foil air journal bearing performance. A known misalignment of increasing severity will be induced to determine the sensitivity of foil bearings to misalignment. Other future plans include oil-free integral starter generator testing and development, and dynamic load testing of foil journal bearings.

Howard, Samuel A.

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING

The ballad of LLM agents: philosophical reasoning for chemistry

Large language models (LLMs) show remarkable potential for scientific reasoning but often produce unreliable or scientifically unactionable outputs when faced with multi-step logic, domain grounding, and interpretability challenges, especially in complex fields like chemistry and materials science. Here, we introduce a framework of philosophical reasoning agents, inspired by canonical thinkers such as Socrates, Descartes, Kant, and Hume, to guide LLM behavior via structured prompt engineering. These agents embody distinct reasoning paradigms (dialectical inquiry, deductive logic, rule-based judgment, and empirical validation) and are evaluated across multiple chemistry subdomains, physical, analytical, general, inorganic, and organic chemistry, using the ChemBench benchmark. Our agentic prompting approach yields substantial accuracy gains on open-ended numerical chemistry questions, with gains of +11.5 percentage points for GPT-4o with Hume, +4.5 percentage points for GPT-5 with Kant, and +21.8 percentage points for GPT-5.1 with Socrates at the strict 1% error threshold, relative to the corresponding base models. Beyond accuracy, we observe benchmark-level model–agent performance patterns, suggesting that different prompting styles interact differently with each base model. These findings demonstrate that embedding philosophy-of-science principles into multi-agent frameworks can improve and produce interpretable, adaptive, and domain-aligned scientific LLMs.

Harb, Hassan [Argonne National Laboratory (ANL), A

Towards philosophical reasoning with agentic LLMs: Socratic method for scientific assistance

As large language models (LLMs) become central tools in science, improving their reasoning capabilities is critical for meaningful and trustworthy applications. We introduce a Socratic agent for scientific reasoning, implemented through a structured system prompt that guides LLMs via classical principles of inquiry. Unlike typical prompt engineering or retrieval-based methods, our approach leverages definition, analogy, hypothesis elimination, and other Socratic techniques to generate more coherent, critical, and domain-aware responses. We evaluate the agent across diverse scientific domains and benchmark it on the abstraction and reasoning corpus challenge dataset, achieving 97.15% under a fixed prompting protocol and without fine-tuning or external tools. Expert evaluation shows improved reasoning depth, clarity, and adaptability over conventional LLM outputs, suggesting that structured prompting rooted in philosophical reasoning can improve the scientific utility of language models.

LLM reasoning

Automatic building energy model development and debugging using large language models agentic workflow

Building energy modeling (BEM) is a complex process that demands significant time and expertise, limiting its broader application in building design and operations. While Large Language Models (LLMs) agentic workflow have facilitated complex engineering processes, their application in BEM has not been specifically explored. This paper investigates the feasibility of automating BEM using LLM agentic workflow. Here, we developed a generic LLM-planning-based workflow that takes a building description as input and generates an error-free EnergyPlus building energy model. Our robust workflow includes four core agents: 1) Building Description Pre-Processing, 2) IDF Object Information Extraction, 3) Single IDF Object Generator Suite, and 4) IDF Debugging Agent. These agents divide the complex tasks into manageable sub-steps, enabling LLMs to generate accurate and reliable results at each stage. The case study demonstrates the successful translation of a building description into an error-free EnergyPlus model for the iUnit modular building at the National Renewable Energy Laboratory. The effectiveness of our workflow surpasses: 1) naive prompt engineering, 2) other LLM-based workflows, and 3) manual modeling, in terms of accuracy, reliability, and time efficiency. The paper concludes with a discussion on the interplay between foundational models and LLM agent planning design, advocating for the use of fine-tuned, specialized models to advance this field.

97 MATHEMATICS AND COMPUTING

A versatile machine learning workflow for high-throughput analysis of supported metal catalyst particles

Accurate and efficient characterization of nanoparticles (NPs), particularly regarding particle size distribution, is essential for advancing our understanding of their structure-property relationship and facilitating their design for various applications. In this study, we introduce a novel two-stage artificial intelligence (AI)-driven workflow for NP analysis that leverages prompt engineering techniques from state-of-the-art single-stage object detection and large-scale vision transformer (ViT) architectures. This methodology is applied to transmission electron microscopy (TEM) and scanning TEM (STEM) images of heterogeneous catalysts, enabling high-resolution, high-throughput analysis of particle size distributions for supported metal catalyst NPs. The model's performance in detecting and segmenting NPs is validated across diverse heterogeneous catalyst systems, including various metals (Ru, Cu, PtCo, and Pt), supports (silica (SiO 2 ), γ-alumina (γ-Al 2 O 3 ), and carbon black), and particle diameter size distributions with mean and standard deviations ranging from 1.6 ± 0.2 nm to 9.7 ± 4.6 nm. The proposed machine learning (ML) methodology achieved an average F1 overlap score of 0.91 ± 0.01 and demonstrated the ability to disentangle overlapping NPs anchored on catalytic support materials. The segmentation accuracy is further validated using the Hausdorff distance and robust Hausdorff distance metrics, with the 90th percent of the robust Hausdorff distance showing errors within 0.4 ± 0.1 nm to 1.4 ± 0.6 nm. In conclusion, our AI-assisted NP analysis workflow demonstrates robust generalization across diverse datasets and can be readily applied to similar NP segmentation tasks without requiring costly model retraining.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

CACTUS: Chemistry Agent Connecting Tool Usage to Science

Large language models (LLMs) have shown remarkable potential in various domains but often lack the ability to access and reason over domain-specific knowledge and tools. In this article, we introduce Chemistry Agent Connecting Tool-Usage to Science (CACTUS), an LLM-based agent that integrates existing cheminformatics tools to enable accurate and advanced reasoning and problem-solving in chemistry and molecular discovery. We evaluate the performance of CACTUS using a diverse set of open-source LLMs, including Gemma-7b, Falcon-7b, MPT-7b, Llama3-8b, and Mistral-7b, on a benchmark of thousands of chemistry questions. Our results demonstrate that CACTUS significantly outperforms baseline LLMs, with the Gemma-7b, Mistral-7b, and Llama3-8b models achieving the highest accuracy regardless of the prompting strategy used. Moreover, we explore the impact of domain-specific prompting and hardware configurations on model performance, highlighting the importance of prompt engineering and the potential for deploying smaller models on consumer-grade hardware without a significant loss in accuracy. By combining the cognitive capabilities of open-source LLMs with widely used domain-specific tools provided by RDKit, CACTUS can assist researchers in tasks such as molecular property prediction, similarity searching, and drug-likeness assessment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

SmileyLlama: modifying large language models for directed chemical space exploration

Here we show that large language models (LLMs) can be transformed via supervised fine-tuning of engineered prompts into SmileyLlama for exploring the chemical space of drug molecules. We benchmark SmileyLlama against pretrained LLMs and chemical language models trained from scratch for generating valid and novel drug-like molecules, and use direct preference optimization to both improve SmileyLlama’s adherence to a prompt and as part of the iMiner reinforcement learning framework to predict molecules with optimized three-dimensional conformations and high binding affinity to drug targets. By training an LLM to speak directly as a chemical language model, while retaining most of its natural language capabilities, we show that SmileyLlama can reliably generate molecules with user-specified properties rather than acting only as a chatbot with knowledge of chemistry or as a virtual assistant. While SmileyLlama is geared toward drug discovery, the supervised fine-tuning/direct preference optimization/LLM framework can be extended to other chemical, biological and materials applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)

Agentic AI vs ML-Based Autotuning: A Comparative Study for Loop Reordering Optimization

High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.

Rosas, Miguel Romero

Prompt Phrase Ordering Using Large Language Models in HPC: Evaluating Prompt Sensitivity

Large language models (LLMs) have demonstrated effective performance in domain-specific tasks, often requiring a well-designed prompt to guide their responses. However, optimizing the right prompt is challenging due to prompt sensitivity—the phenomenon where small changes in the prompt can lead to significant variations in performance. In this study, we evaluate prompt performance by examining all permutations of independent phrases to investigate prompt sensitivity and robustness. We used two datasets: the GSM8k dataset, which assesses mathematical reasoning, and a custom template prompt for summarizing database metadata. Our goal was to evaluate the performance across all permutations of a sequence of prompt phrases. The study was conducted using the llama3-instruct- 7B model hosted on Ollama, with computations parallelized in a high-performance computing environment. By comparing the average index of phrases in the best and worst-performing prompts, we found that the order of independent phrases within a prompt significantly impacts LLM performance. Additionally, we used Hamming distance to assess changes between phrase orderings, concluding that prompt modifications can dramatically affect scores, often by almost random chance. These findings support existing research on prompt sensitivity. We discuss the challenges of prompt optimization, noting that altering phrases in a successful prompt does not always result in another successful prompt.

97 MATHEMATICS AND COMPUTING

The Mars Climate Sounder In-Flight Positioning Anomaly

The paper discusses the Mars Climate Sounder (MCS) instrument s in-flight positioning errors and presents background material about it. A short overview of the instrument s science objectives and data acquisition techniques is provided. The brief mechanical description familiarizes the reader with the MCS instrument. Several key items of the flight qualification program, which had a rigorous joint drive test program but some limitations in overall system testing, are discussed. Implications this might have had for the flight anomaly, which began after several months of flawless space operation, are mentioned. The detection, interpretation, and instrument response to the errors is discussed. The anomaly prompted engineering reviews, renewed ground, and some in-flight testing. A summary of these events, including a timeline, is included. Several items of concern were uncovered during the anomaly investigation, the root cause, however, was never found. The instrument is now used with two operational constraints that work around the anomaly. It continues science gathering at an only slightly diminished pace that will yield approximately 90% of the originally intended science.

Jau, Bruno M.