Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Prediction algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

An Advanced Machine Learning and Artificial Intelligence System for Demonstrating Radiation Regulatory Compliance in DOE Accelerator Facilities

In this Phase II proposal, Applied Research LLC (ARLLC), Thomas Jefferson National Accelerator Facility (Jefferson Lab), and Old Dominion University (ODU) propose the combination of domain knowledge (beam characteristics, fixed structural shielding, earthen burden (the soil and foliage added to the dome of the experimental halls as additional shielding), etc.), machine learning (ML) and/or artificial intelligence (AI) to correlate a variety of multi-modal onsite signals and the radiation fields seen in accessible areas of the accelerator site and the site boundary. The ML/AI will consider the complex influence of environmental parameters affecting the radon contribution of the measurements, focusing on actual data obtained from Jefferson Lab. In Phase I, the coded beam and location data were fed into a deep learning model to predict doses at several designated locations in Jefferson Lab’s facility. Moreover, a dense radiation map was generated using only a sparse collection of the samples in a facility. In Phase II, we will develop a software prototype containing a radiation prediction algorithm, dense radiation map algorithms, and background noise prediction algorithms, with actual data used to evaluate the prototype. This work will provide a framework for evaluation of radiation measurement results around the site based on learned responses. In addition, the proposed approach allows more granular mapping of radiation levels. Better understanding and communication of these levels is related to the overall approach in keeping doses to personnel ALARA.

43 PARTICLE ACCELERATORS

Inference of response functions with the help of machine-learning algorithms

Response functions are a key quantity to describe the near-equilibrium dynamics of strongly interacting many-body systems. Recent techniques that attempt to overcome the challenges of calculating these ab initio have employed expansions in terms of orthogonal polynomials. We employ a neural network prediction algorithm to reconstruct a response function 𝑆⁡(𝜔) defined over a range in frequencies 𝜔. Here, we represent the calculated response function as a truncated Chebyshev series whose coefficients can be optimized to reduce the representation error. We compare the quality of response functions obtained using coefficients calculated using a neural network (NN) algorithm with those computed using the Gaussian integral transform (GIT) method. In the regime where only a small number of terms in the Chebyshev series are retained, we find that the NN scheme outperforms the GIT method.

Kurkcuoglu, Doga Murat [Fermi National Accelerator

Exploring the Model Design Space for Battery Health Management

Battery Health Management (BHM) is a core enabling technology for the success and widespread adoption of the emerging electric vehicles of today. Although battery chemistries have been studied in detail in literature, an accurate run-time battery life prediction algorithm has eluded us. Current reliability-based techniques are insufficient to manage the use of such batteries when they are an active power source with frequently varying loads in uncertain environments. The amount of usable charge of a battery for a given discharge profile is not only dependent on the starting state-of-charge (SOC), but also other factors like battery health and the discharge or load profile imposed. This paper presents a Particle Filter (PF) based BHM framework with plug-and-play modules for battery models and uncertainty management. The batteries are modeled at three different levels of granularity with associated uncertainty distributions, encoding the basic electrochemical processes of a Lithium-polymer battery. The effects of different choices in the model design space are explored in the context of prediction performance in an electric unmanned aerial vehicle (UAV) application with emulated flight profiles.

Saha, Bhaskar

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING

Structure-aware Initialization via Numerical Continuation and Informed Priors

Scientific machine learning (SciML) often operates in ill-conditioned, weakly identifiable regimes due to limited data or indirect observations. In such settings, optimization and inference are highly sensitive to the starting point, making initialization--often under-reported--a consequential degree of freedom. Random initialization is not a neutral default as it induces an implicit prior over candidate solutions and can systematically bias the result, producing large run-to-run variability. Here, we formalize this view by treating initialization as a hidden confounder in SciML and develop a unifying theory for structure-aware initialization via numerical continuation, constructing warm starts from related problem instances. Across representative tasks, including physics-informed neural networks, maximum likelihood estimation, and variational inference, warm starts have been shown to consistently reduce optimization effort and improve reliability.

Data integrity

Optimization of an aerostructural machining process using physics-guided Bayesian stability modelling

Existing algorithms for predicting milling chatter have not been widely adopted in industry since they require specialized instruments to measure the stability inputs. This study describes how the machining process for a meter-scale aluminum aerostructure was optimized using a physics-guided Bayesian stability model. The study was performed in collaboration with an industrial partner on production machines to evaluate the practicality of the proposed method under real-world conditions. For each cutting tool, the Bayesian approach automatically selected a small number of cutting tests, which were monitored using a microphone to observe the chatter frequency. The algorithm learned the system dynamics, cutting forces, and stability map from these test results. A novel algorithm for predicting tool bending stress was incorporated into the test selection algorithm to avoid tool breakage. On average, each set of optimized cutting parameters required less than six tests to identify and were 97% more productive than baseline parameters from the cutting tool manufacturer. The machining program was then further optimized using commercial feedrate scheduling software to remove cutting force spikes and reduce air cutting time. Five components were machined using the optimized process. These results demonstrate the potential for physics-guided Bayesian models to improve productivity in industrial settings.

Cornelius, Aaron [UT Knoxville]

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology

Feature engineering descriptors, transforms, and machine learning for grain boundaries and variable-sized atom clusters

Abstract Obtaining microscopic structure-property relationships for grain boundaries is challenging due to their complex atomic structures. Recent efforts use machine learning to derive these relationships, but the way the atomic grain boundary structure is represented can have a significant impact on the predictions. Key steps for property prediction common to grain boundaries and other variable-sized atom clustered structures include: (1) describing the atomic structure as a feature matrix, (2) transforming the variable-sized feature matrix to a fixed length common to all structures, and (3) applying a machine learning algorithm to predict properties from the transformed matrices. We examine how these steps and different combinations of engineered features impact the accuracy of grain boundary energy predictions using a database of over 7000 grain boundaries. Additionally, we assess how different engineered features support interpretability, offering insights into the physics of the structure-property relationships.

36 MATERIALS SCIENCE

Enabling integrated AI control on DIII-D: a control system design with state-of-the-art experiments

We present the design and application of a general algorithm for Prediction And Control using MAchiNe learning (PACMAN) in DIII-D. Machine learning (ML)-based predictors and controllers have shown great promise in achieving regimes in which traditional controllers fail, such as tearing mode (TM) free scenarios, ELM-free scenarios and stable advanced tokamak conditions. The architecture presented here was deployed on DIII-D to facilitate the end-to-end implementation of advanced control experiments, from diagnostic processing to final actuation commands. This paper describes the detailed design of the algorithm and explains the motivation behind each design point. We also describe several successful ML control experiments in DIII-D using this algorithm, including a reinforcement learning controller targeting advanced non-inductive plasmas, a wide-pedestal quiescent H-mode ELM predictor, an Alfvén Eigenmode controller, a Model Predictive Control plasma profile controller and a state-machine TM predictor-controller. There is also discussion on guiding principles for real-time ML controller design and implementation.

machine learning

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka

Extending Component Lifetime And Improving Inverter Reliability (ECLAIIR)

Inverter reliability remains one of the most persistent challenges limiting the performance, availability, and economic viability of utility‑scale photovoltaic (PV) plants. Industry data consistently show that inverters account for the highest share of corrective maintenance events and unplanned outages across PV fleets. These failures result in energy losses, increased O&M costs, and reduced confidence in long‑term solar asset performance. Motivated by these challenges, this project—Extending Component Lifetime and Improving Inverter Reliability (ECLAIIR)—was undertaken to systematically investigate inverter degradation and failure mechanisms, develop predictive maintenance capabilities, and establish data‑driven pathways to improve service life and reduce the Levelized Cost of Energy (LCOE) for large‑scale PV systems. The primary goal of the project was to identify pre‑failure signatures in string inverters using both lab‑based accelerated lifetime testing and field‑based data and to develop predictive maintenance algorithms that can anticipate inverter faults before they occur. Through collaboration with inverter testing laboratory, solar PV plant owner, and failure‑analysis experts, the project advanced the technical understanding of inverter reliability. By instrumenting inverters with thermistors, humidity sensors, power‑quality meters, and acoustic sensors, the research established how multiple sensing modalities can reliably detect deviations from normal behavior hours to days before failure. These findings substantially enhance scientific understanding of inverter failure kinetics and provide the PV industry with the most comprehensive cross‑OEM characterization of early‑stage failure indicators reported to date. Technically, the project demonstrated the effectiveness of predictive maintenance by developing and validating the PreDICT (Predictive Diagnostics of PV Inverters Using Condition Monitoring and Trend Analysis) framework—a multi‑layer diagnostic architecture combining peer‑to‑peer analytics, historical trend modeling, and advanced machine‑learning techniques such as the Sequential Conditional Variational Autoencoder (SCVAE). This predictive model achieved more than 90% accuracy in detecting pre‑failure conditions and provided up to four days of lead time before inverter failure in field scenarios. Economically, the project’s LCOE analysis showed that predictive maintenance can reduce lifetime energy losses and minimize corrective maintenance interventions. Modeling indicated that, depending on inverter failure rates and replacement timelines, predictive maintenance can significantly reduce LCOE impacts associated with inverter downtime: from as high as 19.4% under conventional maintenance strategies to 0.1%–10.17% when predictive analytics are adopted. These results confirm that predictive maintenance is both technically feasible and economically advantageous for utilities and plant operators. The project’s findings also have broad public benefit. By improving inverter reliability and reducing downtime, predictive maintenance directly increases electricity generation from existing PV assets. Enhanced reliability lowers operational costs for utilities, which can translate over time into lower energy costs for consumers. Furthermore, the project’s technical publications, conference presentations, and industry workshops ensure that knowledge gained is shared broadly across the solar industry, supporting workforce development and enabling utilities of all sizes to adopt modern asset‑health monitoring practices. The retrofitting case study and service‑life prediction framework further support informed decision‑making for aging PV fleets, helping operators extend system life and reduce electronic waste. In summary, the ECLAIIR project significantly advanced the state of knowledge on inverter degradation, demonstrated the technical and economic value of predictive maintenance, and delivered actionable tools and insights that support more reliable, cost‑effective, and sustainable PV plant operation. The outcomes of this project will continue to inform utility practices, guide inverter design improvements, and strengthen the long‑term performance of solar assets nationwide.

14 SOLAR ENERGY

Models and Algorithms for Equilibrium Analysis of Mixed-Material Nucleic Acid Systems

Dynamic programming algorithms within the NUPACK software suite enable analysis of equilibrium base-pairing properties for complex and test tube ensembles containing arbitrary numbers of interacting nucleic acid strands. Currently, calculations are limited to single-material systems that are either all-RNA or all-DNA. Here, to enable analysis of mixed-material systems that are critical for modern applications in vitro, in situ, and in vivo, we develop physical models and dynamic programming algorithms that allow the material of the system to be specified at nucleotide resolution. Free energy parameter sets are constructed for both RNA/DNA and RNA/2'OMe-RNA mixed-material systems by combining available empirical mixed-material parameters with single-material parameter sets to enable treatment of the full complex and test tube ensembles. New dynamic programming recursions account for the material of each nucleotide throughout the recursive process. For a complex with N nucleotides, the mixed-material dynamic programming algorithms maintain the O(N 3 ) time complexity of the single-material algorithms, enabling efficient calculation of diverse physical quantities over complex and test tube ensembles (e.g., complex partition function, equilibrium complex concentrations, equilibrium base-pairing probabilities, minimum free energy secondary structure(s), and Boltzmann-sampled secondary structures) at a cost increase of roughly 2.0-3.5×. The results of existing single-material algorithms are exactly reproduced when applying the new mixed-material algorithms to single-material systems. Accuracy is significantly enhanced using mixed-material models and algorithms to predict RNA/DNA and RNA/2'OMe-RNA duplex melting temperatures from the experimental literature as well as RNA/DNA melt profiles from new experiments. In conclusion, mixed-material analyses can be performed online using the NUPACK web app (www.nupack.org) or locally using the NUPACK Python module.

2′OMe-RNA

Trajectory Shaper: A Solution for Disrupted Cooperative Adaptive Cruise Control

Cooperative adaptive cruise control (CACC) can effectively reduce energy consumption, alleviate traffic congestion, and enhance safety. However, communication-related constraints and uncooperative vehicle users can disrupt CACC during real-world operations, significantly undermining the putative benefits of CACC. To alleviate the negative impacts of disrupted CACC, this study develops the trajectory shaper (TS) methods as backup solutions for two scenarios: (i) communication between vehicles is infeasible, and vehicles execute adaptive cruise control (ACC) using local sensor measurements; (ii) follower vehicles reject forming a cooperative platoon and execute their local distributed controllers using the information attained via communication. When communication is infeasible, a distributed TS is devised on each vehicle to modify the sensor measurements, enabling safe and efficient ACC operations. When communication is available but uncooperative agents are involved, the lead vehicle of the platoon executes a centralized TS to modify the information shared with uncooperative agents, achieving optimal platoon-level performance. The centralized and distributed TSs are implemented based on the model predictive control algorithms to yield optimal modifications on input information. Robustness is also factored to tackle model uncertainties during TS operations to ensure safety and efficiency. Numerical experiments validate the control performance of the proposed TSs.

Zhou, Anye [ORNL] (ORCID:0000000301455579)

Stochastic Model Predictive Control With Gaussian Wind Direction Preview for Wake Steering

This article addresses the problem of wake steering control for wind farms that explicitly consider the tradeoff between farm-level power generation and yaw duty cycle under variable and uncertain wind conditions. A novel stochastic model predictive control (MPC) algorithm is presented, which utilizes a stochastic model of the freestream wind field components in a receding horizon framework to compute optimal yaw set points that maximize the expected value of the farm power while constraining the yaw actuation. Different configurations of the algorithm are evaluated using a steady-state wind farm simulator. The proposed stochastic MPC algorithm can plan control actions over a future prediction horizon based on probabilistic estimates of the incoming wind magnitude and direction.

17 WIND ENERGY

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy