Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AI efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Understanding the High Energy Higgs Sector with the CMS Experiment and Artificial Intelligence

This dissertation describes efforts towards understanding the Higgs boson at the highest energies humanly accessible, using the CMS experiment at the Large Hadron Collider and advances in artificial intelligence (AI) and machine learning (ML). We present searches for resonant and nonresonant Higgs-boson (H) pair production in the all-hadronic two beauty-quark and two vector boson (V) final state, using a novel strategy to measure the quartic HHVV coupling and search for new Higgs-like bosons. By targeting highly Lorentz-boosted Higgs pairs, we probe effects of potential new physics in the high energy Higgs sector, which could hold answers to fundamental mysteries of nature such as baryon asymmetry. To enable these and future searches, we introduce as well significant developments in AI/ML, including in the identification of boosted H$\rightarrow$VV decays with deep transformer networks and advances in AI-accelerated fast simulations of the CMS detector. The latter notably includes the development of the first, highly performant generative models for point-cloud data in high energy physics, which have the potential to improve CMS' computational efficiency by up to three orders of magnitude. We also highlight novel solutions to the important and challenging problems of calibrating and validating these ML techniques. Finally, we present new approaches to search for new physics in a model-agnostic manner, using physics-informed ML methods equivariant to Lorentz transformations. The quartic HHVV coupling is observed (expected) to be constrained to $[-0.04, 2.05]$ ($[0.05, 1.98]$) at the 95% confidence level relative to the standard model prediction, representing the second-most sensitive measurement of this coupling by CMS to date. Exclusion limits on the production cross section of new heavy resonances decaying to two Higgs-like bosons are expected to be as low as 0.3 fb for high resonance masses.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Grand Challenge "Uncertainty Project" to Accelerate Advances in Earth System Predictability: AI-Enabled Concepts and Applications

This proposal is emerging from GISS ModelE3 ESM development in the area of cloud physics, so we begin with an example of research needs/gaps from that work. Here, some of our greatest development concerns arise where we lack fundamental process-level understanding, as in ice formation. Namely, it is currently unclear what is the main process that is forming the majority of ice crystals in commonly occurring convection, apparently via secondary ice production at warm temperatures. We are keenly awaiting laboratory data for candidate mechanisms, which is not yet in hand to crucially establish their efficiency. Our progress is also hampered by a lack of uncertainty characterization in currently available measurements of ice crystal number size distributions. Furthermore, the same multiplication process may be responsible for a majority of ice crystals in many extratropical mixed-phase clouds, whose variable representation in CMIP6 ESMs may be a leading cause of differences in cloud phase feedback and ECS. Yet we have been required to deliver an ESM with the cloud physics knowledge at hand. The proposed grand challenge project is AI-enabled via application of machine learning (ML) to climate model and observational data streams (focal area 3), and applications include AI-guided observing system design and model/component/parameterization selection (areas 1 and 2). The project is structurally agnostic as to whether model or observing system components use AI approaches or not, but uncertainties must be estimated and propagatable in both.

58 GEOSCIENCES↗

The Next Stage for Smart Cities: Metrics to Spur an Evolution from Technology for Efficiency to Technology for People's Needs

The Smart Cities movement is about a decade old, energized by the USDOT Smart Cities grant that sparked cities to dream big. Now, 10 years later, and multitudes of conferences and papers later, the biggest question regarding Smart Cities is whether any progress has been made? Smart Cities forums are typically led by the technology sector, either telecommunications, internet access, or most recently artificial intelligence (AI) concerns. Each of these areas have grown and found new applications in the context of Smart Cities, however, these technology applications are just tools toward ends. Do digital communications and assistance enable or foster improved access to real, tactile, non-virtual goods and services such as transportation, food, housing, trash removal, health care, education, social interaction? These and other topics are not 'big tech', but are each critical to quality of life outcomes and are day-to-day necessities, comprising many of the touch points through which citizens interact with municipal agencies. Regarding Smart Cities, relationships have not been fully refined as to how best to merge the emerging generation of technology and related trends with more mundane but essential requirements of tactile services. How should municipal managers and operators bring the range of new technology tools (e.g. on-demand transit, automated vehicles, digital connectivity) to bear to improve functions serving real- physical- non-virtual needs? How do we make cities more efficient in providing access to services and opportunities for improved quality of life? This paper examines the current Smart City movement, emerging trends, and the need to put metrics with respect to the delivery, access, and provision of real-life needs. Using a few case studies of past, present, and projected future scenarios, this paper attempts to bring Smart Cities into the real cities domain, perhaps pushing the moniker away from 'Smart' to 'Efficient' - something that can be measured, counted, and assessed. Through the consideration of the next stage of the evolution of cities from 'Smart' to 'Efficient', the combined efforts and experience of multiple researchers at the NREL and their network of collaborators highlights the need for updates metrics that can be used to start guiding efforts toward productive initiatives to improve quality of life.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

Hybrid Recurrent Neural Network Modeling for Traffic Delay Prediction at Signalized Intersections Along an Urban Arterial

This paper studies the traffic delay prediction modeling for multiple signalized intersections along the Ala Moana Boulevard and Nimitz Highway in Hawaii. Several machine learning (ML) based approaches have been studied in the literature, and most of them focused on prediction accuracy rather than the end use of real-time control and implementation. These ML models tend to be very complex and non-linear in nature, making it challenging to achieve fast inferences and are computationally heavy for real-time signal control implementation. As such, in this paper, a simple yet accurate hybrid modeling method is proposed to predict traffic delay one-step ahead with the model made suitable for real-time implementation to control traffic flow. Since real-time road-side measurements are recorded in unstructured form, the paper also discusses other issues related to data extraction and the pre-processing process. Finally, a simple signal control loop is developed to demonstrate the proposed modeling approach, which has shown advantages in model accuracy and computation efficiency compared against several existing modeling methods.

42 ENGINEERING↗

AI-assisted detector design for the EIC (AID(2)E)

Artificial Intelligence is poised to transform the design of complex, large-scale detectors like ePIC at the future Electron Ion Collider. Featuring a central detector with additional detecting systems in the far forward and far backward regions, the ePIC experiment incorporates numerous design parameters and objectives, including performance, physics reach, and cost, constrained by mechanical and geometric limits. This project aims to develop a scalable, distributed AI-assisted detector design for the EIC (AID(2)E), employing state-of-the-art multiobjective optimization to tackle complex designs. Supported by the ePIC software stack and using G EANT 4 simulations, our approach benefits from transparent parameterization and advanced AI features. The workflow leverages the PanDA and iDDS systems, used in major experiments such as ATLAS at CERN LHC, the Rubin Observatory, and sPHENIX at RHIC, to manage the compute intensive demands of ePIC detector simulations. Tailored enhancements to the PanDA system focus on usability, scalability, automation, and monitoring. Ultimately, this project aims to establish a robust design capability, apply a distributed AI-assisted workflow to the ePIC detector, and extend its applications to the design of the second detector (Detector-2) in the EIC, as well as to calibration and alignment tasks. Additionally, we are developing advanced data science tools to efficiently navigate the complex, multidimensional trade-offs identified through this optimization process.

97 MATHEMATICS AND COMPUTING↗

Heuristic Evaluation Methods Applied to a Predictive Maintenance Chatbot

The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER’s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use.

99 - GENERAL AND MISCELLANEOUS↗

Self-testing of a single quantum system from theory to experiment

Abstract Self-testing allows one to characterise quantum systems under minimal assumptions. However, existing schemes rely on quantum nonlocality and cannot be applied to systems that are not entangled. Here, we introduce a robust method that achieves self-testing of individual systems by taking advantage of contextuality. The scheme is based on the simplest contextuality witness for the simplest contextual quantum system—the Klyachko-Can-Binicioğlu-Shumovsky inequality for the qutrit. We establish a lower bound on the fidelity of the state and the measurements as a function of the value of the witness under a pragmatic assumption on the measurements. We apply the method in an experiment on a single trapped 40 Ca + using randomly chosen measurements and perfect detection efficiency. Using the observed statistics, we obtain an experimental demonstration of self-testing of a single quantum system.

Physics↗

EnergyPlus-MCP: A model-context-protocol server for ai-driven building energy modeling

Traditional building energy modeling with the EnergyPlus building performance simulation engine requires domain expertise, programming skills, and intensive manual efforts limiting its effective adoption. This paper introduces EnergyPlus-MCP, the first open-source Model Context Protocol (MCP) server specifically designed for EnergyPlus simulation workflows, establishing a new foundational infrastructure for AI-driven building energy modeling. The MCP server implements a layered architecture with 35 specialized tools spanning model management, editing and analysis, HVAC and other systems configuration inspection, and simulation execution, enabling Large Language Models to interact with EnergyPlus through conversational interfaces. The server addresses critical workflow barriers by automating model validation, streamlining energy efficiency measures modification, and providing intelligent output management with interactive visualization. Through practical demonstrations using a multi-zone building retrofit analysis, we show how the EnergyPlus-MCP server significantly reduces manual efforts while maintaining full simulation rigor. By providing accessible natural language interfaces to sophisticated building energy analysis, this approach enables scalable deployment of simulation expertise across public and private organizations, educational institutions, and research teams, fundamentally transforming traditional building energy modeling practices.

AI↗

Teaching Freight Mode Choice Models New Tricks Using Interpretable Machine Learning Methods

Understanding and forecasting the intricate freight mode choice behavior under various industry, policy, and technology contexts is essential in freight planning and policymaking. Numerous models have been developed in prior studies to provide insights into freight mode selection, the majority of which use discrete choice models such as multinomial logit (MNL) models. However, logit models often rely on linear specifications of independent variables, despite potential nonlinear relationships in the data. Moreover, there often lacks a heuristic and efficient approach to identify such complex relationships to define the logit model specifications. To fill this gap, we developed an MNL model for freight mode choice using the insights from state-of-the- art machine learning (ML) models. ML models can capture the nonlinear nature of the complex decision-making process, and recent advances in 'explainable AI' have greatly improved their interpretability. The interpretable ML methods help enhance the performance of MNL models and advance knowledge of freight mode choice. Specifically, the influential factors and their relationship with individual modes are identified using SHapley Additive exPlanations (SHAP) to improve the MNL's performance. The workflow is demonstrated in a case study of Austin, Texas, and the SHAP results reveal multiple nonlinear relationships predicted by ML models. Incorporating those relationships into MNL model specifications improves the interpretability and accuracy of the MNL model compared to a conventional MNL model. Findings from this study can be used to guide freight planning and inform policymakers and practitioners on how key factors affect freight decision-making.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

A versatile machine learning workflow for high-throughput analysis of supported metal catalyst particles

Accurate and efficient characterization of nanoparticles (NPs), particularly regarding particle size distribution, is essential for advancing our understanding of their structure-property relationship and facilitating their design for various applications. In this study, we introduce a novel two-stage artificial intelligence (AI)-driven workflow for NP analysis that leverages prompt engineering techniques from state-of-the-art single-stage object detection and large-scale vision transformer (ViT) architectures. This methodology is applied to transmission electron microscopy (TEM) and scanning TEM (STEM) images of heterogeneous catalysts, enabling high-resolution, high-throughput analysis of particle size distributions for supported metal catalyst NPs. The model's performance in detecting and segmenting NPs is validated across diverse heterogeneous catalyst systems, including various metals (Ru, Cu, PtCo, and Pt), supports (silica (SiO 2 ), γ-alumina (γ-Al 2 O 3 ), and carbon black), and particle diameter size distributions with mean and standard deviations ranging from 1.6 ± 0.2 nm to 9.7 ± 4.6 nm. The proposed machine learning (ML) methodology achieved an average F1 overlap score of 0.91 ± 0.01 and demonstrated the ability to disentangle overlapping NPs anchored on catalytic support materials. The segmentation accuracy is further validated using the Hausdorff distance and robust Hausdorff distance metrics, with the 90th percent of the robust Hausdorff distance showing errors within 0.4 ± 0.1 nm to 1.4 ± 0.6 nm. In conclusion, our AI-assisted NP analysis workflow demonstrates robust generalization across diverse datasets and can be readily applied to similar NP segmentation tasks without requiring costly model retraining.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

VISIONARY: Virtual Intelligence System for Optimizing Novel Analytical Research Yields

VISIONARY is an AI system that accelerates energy materials discovery by automatically generating hypotheses about structure-property relationships. It analyzes patterns in materials data, identifies promising correlations, and proposes testable scientific hypotheses without human intervention. By streamlining this reasoning process, VISIONARY helps researchers efficiently identify candidate materials with desired properties, significantly speeding up the materials development pipeline for energy applications. During the project, we developed a standalone application. The application uses a combination of papers provided by the user and data collected from FutureHouse’s dataset to build an understanding of the background that the user wants to explore for the hypothesis.

36 MATERIALS SCIENCE↗

Open-source FPGA-ML codesign for the MLPerf Tiny Benchmark

We present our development experience and recent results for the MLPerf Tiny Inference Benchmark on field-programmable gate array (FPGA) platforms. We use the open-source hls4ml and FINN workflows, which aim to democratize AI-hardware codesign of optimized neural networks on FPGAs. We present the design and implementation process for the keyword spotting, anomaly detection, and image classification benchmark tasks. The resulting hardware implementations are quantized, configurable, spatial dataflow architectures tailored for speed and efficiency and introduce new generic optimizations and common workflows developed as a part of this work. The full workflow is presented from quantization-aware training to FPGA implementation. The solutions are deployed on system-on-chip (Pynq-Z2) and pure FPGA (Arty A7-100T) platforms. The resulting submissions achieve latencies as low as 20 $\mu$s and energy consumption as low as 30 $\mu$J per inference. We demonstrate how emerging ML benchmarks on heterogeneous hardware platforms can catalyze collaboration and the development of new techniques and more accessible tools.

Borras, Hendrik↗

Li-ion battery design through microstructural optimization using generative AI

Lithium-ion batteries are used across various applications, necessitating tailored cell designs to enhance performance. Optimizing electrode manufacturing parameters is a key route to achieving this, as these parameters directly influence the microstructure and performance of the cells. However, linking process parameters to performance is complex, and experimental or modeling campaigns are often slow and expensive. This study introduces a fast computational optimization framework for electrode manufacturing parameters. A generative model, trained on a small dataset of microstructural images associated with different manufacturing parameters, efficiently generates representative microstructures for new parameters. This model is integrated into a Bayesian optimization loop that includes microstructure generation, characterization, and simulation, aiming to find optimal manufacturing parameters for a particular application. Significant improvement in the energy density of a 4680 cell is achieved through bespoke cell design, highlighting the importance of cell-scale normalization. The framework’s modularity allows its application to various advanced materials manufacturing scenarios.

batteries↗

Transforming ENERGY through Computational Excellence

Computational methods underpin advancing the science and engineering of energy efficiency, sustainable transportation, renewable power technologies, and developing a knowledge base to optimize energy systems. Researchers with access to enough computing, and the right type, can focus their ingenuity and creativity on addressing the energy challenges. NREL’s advanced computing influence spans several common themes across the Office of Energy Efficiency and Renewable Energy (EERE), including materials discovery, process modeling, fluid dynamics, resource mapping, and analysis of large-scale systems with real-time optimization.

advanced computing↗

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]↗