Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Prediction Accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Comparison of Preconstruction and Operational Wake Loss Estimates for Land-Based Wind Plants

Recent studies suggest that biases between wind plant pre-construction energy yield estimates and actual energy production are decreasing over time. However, variability in energy yield prediction accuracy across different projects and wind energy consultants remains high. Wake effects are one of the largest categories comprising the pre-construction energy yield assessment process. To assess the accuracy of wake loss predictions, we compare pre-construction wake loss estimates provided by 8 consultants to the estimated operational wake losses for 10 North American wind plants, as part of the Wind Plant Performance Prediction (WP3) Benchmark project. We estimate operational wake losses using supervisory control and data acquisition (SCADA) data by comparing total wind plant energy production to the potential energy production based on the power produced by freestream wind turbines. In the presentation, we will discuss the overall wake loss prediction bias as well as the project-to-project variability in the prediction accuracy. Further, we will highlight challenges encountered when estimating operational wake losses, including the impact of complex terrain and the presence of neighboring wind plants.

benchmark↗

A Final Approach Trajectory Model for Current Operations

Predicting accurate trajectories with limited intent information is a challenge faced by air traffic management decision support tools in operation today. One such tool is the FAA's Terminal Proximity Alert system which is intended to assist controllers in maintaining safe separation of arrival aircraft during final approach. In an effort to improve the performance of such tools, two final approach trajectory models are proposed; one based on polynomial interpolation, the other on the Fourier transform. These models were tested against actual traffic data and used to study effects of the key final approach trajectory modeling parameters of wind, aircraft type, and weight class, on trajectory prediction accuracy. Using only the limited intent data available to today's ATM system, both the polynomial interpolation and Fourier transform models showed improved trajectory prediction accuracy over a baseline dead reckoning model. Analysis of actual arrival traffic showed that this improved trajectory prediction accuracy leads to improved inter-arrival separation prediction accuracy for longer look ahead times. The difference in mean inter-arrival separation prediction error between the Fourier transform and dead reckoning models was 0.2 nmi for a look ahead time of 120 sec, a 33 percent improvement, with a corresponding 32 percent improvement in standard deviation.

Gong, Chester↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

Data-driven chaos indicator for nonlinear dynamics and applications on storage ring lattice design

A data-driven chaos indicator concept is introduced to characterize the degree of chaos for nonlinear dynamical systems. The indicator is represented by the prediction accuracy of surrogate models established purely from data. It provides a metric for the predictability of nonlinear motions in a given system. When using the indicator to implement a tune-scan for a quadratic Hénon map, the main resonances and their asymmetric stop-band widths can be identified. When applied to particle transportation in a storage ring, as particle motion becomes more chaotic, its surrogate model prediction accuracy decreases correspondingly. So, the prediction accuracy, acting as a chaos indicator, can be used directly as the objective for nonlinear beam dynamics optimization. This method provides a different perspective on nonlinear beam dynamics and an efficient method for nonlinear lattice optimization. Applications in dynamic aperture optimization are demonstrated as real world examples.

36 MATERIALS SCIENCE↗

Data-driven Chaos Indicator for Nonlinear Dynamics and Applications on Storage Ring Lattice Design

A data-driven chaos indicator concept is introduced to characterize the degree of chaos for nonlinear dynamical systems. The indicator is represented by the prediction accuracy of surrogate models established purely from data. It provides a metric for the predictability of nonlinear motions in a given system. When using the indicator to implement a tune-scan for a quadratic Hénon map, the main resonances and their asymmetric stop-band widths can be identified. When applied to particle transportation in a storage ring, as particle motion becomes more chaotic, its surrogate model prediction accuracy decreases correspondingly. Therefore, the prediction accuracy, acting as a chaos indicator, can be used directly as the objective for nonlinear beam dynamics optimization. This method provides a different perspective on nonlinear beam dynamics and an efficient method for nonlinear lattice optimization. Applications in dynamic aperture optimization are demonstrated as real world examples.

43 PARTICLE ACCELERATORS↗

Training Population Optimization for Genomic Selection in Miscanthus

Miscanthus is a perennial grass with potential for lignocellulosic ethanol production. To ensure its utility for this purpose, breeding efforts should focus on increasing genetic diversity of the nothospecies Miscanthus × giganteus (M×g) beyond the single clone used in many programs. Germplasm from the corresponding parental species M. sinensis (Msi) and M. sacchariflorus (Msa) could theoretically be used as training sets for genomic prediction of M×g clones with optimal genomic estimated breeding values for biofuel traits. To this end, we first showed that subpopulation structure makes a substantial contribution to the genomic selection (GS) prediction accuracies within a 538-member diversity panel of predominately Msi individuals and a 598-member diversity panels of Msa individuals. We then assessed the ability of these two diversity panels to train GS models that predict breeding values in an interspecific diploid 216-member M×g F2 panel. Low and negative prediction accuracies were observed when various subsets of the two diversity panels were used to train these GS models. To overcome the drawback of having only one interspecific M×g F2 panel available, we also evaluated prediction accuracies for traits simulated in 50 simulated interspecific M×g F2 panels derived from different sets of Msi and diploid Msa parents. The results revealed that genetic architectures with common causal mutations across Msi and Msa yielded the highest prediction accuracies. Ultimately, these results suggest that the ideal training set should contain the same causal mutations segregating within interspecific M×g populations, and thus efforts should be undertaken to ensure that individuals in the training and validation sets are as closely related as possible.

59 BASIC BIOLOGICAL SCIENCES↗

A Neural Network Approach to Predict Gibbs Free Energy of Ternary Solid Solutions

Here, we present a data-centric deep learning (DL) approach using neural networks (NNs) to predict the thermodynamics of ternary solid solutions. We explore how NNs can be trained with a dataset of Gibbs free energies computed from a CALPHAD database to predict ternary systems as a function of composition and temperature. We have chosen the energetics of the FCC solid solution phase in 226 binaries consisting of 23 elements at 11 different temperatures to demonstrate the feasibility. The number of binary data points included in the present study is 102,000. We select six ternaries to augment the binary dataset to investigate their influence on the NN prediction accuracy. We examine the sensitivity of data sampling on the prediction accuracy of NNs over selected ternary systems. It is anticipated that the current DL workflow can be further elevated by integrating advanced descriptors beyond the elemental composition and more curated training datasets to improve prediction accuracy and applicability.

42 ENGINEERING↗

ScaleML: Machine Learning based Heap Memory Object Scaling Prediction

Memory subsystem contributes 28-40% of total energy consumption. Several studies investigated energy prediction and consumption via profiling memory object access patterns. However, such profiling leads to higher energy consumption due to intense memory object-level profiling to achieve high prediction accuracy. Further, memory object access pattern prediction has been considered through analyzing the variation between memory object access patterns, referred to as scaling rate. The existing techniques for scaling rate prediction, such as Linear Scaling Rate (LSR), suffer from a high error rate in prediction with changes in access patterns, which leads to a high error rate of energy consumption prediction. In this paper, we compare and evaluate several memory object access pattern prediction models including LSR and machine learning (ML) models. Further, we propose SCALEML, a heap memory object scaling rate prediction mechanism that employs an ML model to achieve high access pattern prediction accuracy with variations in memory object access patterns. We evaluate SCALEML using various application benchmarks. The experimental results show that SCALEML achieves about 20% higher accuracy than the LSR model for predicting the scaling rate of object access patterns and energy estimation.

Park, Joongeon↗

Prediction of histone post-translational modifications using deep learning

Abstract Motivation Histone post-translational modifications (PTMs) are involved in a variety of essential regulatory processes in the cell, including transcription control. Recent studies have shown that histone PTMs can be accurately predicted from the knowledge of transcription factor binding or DNase hypersensitivity data. Similarly, it has been shown that one can predict PTMs from the underlying DNA primary sequence. Results In this study, we introduce a deep learning architecture called DeepPTM for predicting histone PTMs from transcription factor binding data and the primary DNA sequence. Extensive experimental results show that our deep learning model outperforms the prediction accuracy of the model proposed in Benveniste et al. (PNAS 2014) and DeepHistone (BMC Genomics 2019). The competitive advantage of our framework lies in the synergistic use of deep learning combined with an effective pre-processing step. Our classification framework has also enabled the discovery that the knowledge of a small subset of transcription factors (which are histone-PTM and cell-type-specific) can provide almost the same prediction accuracy that can be obtained using all the transcription factors data. Availabilityand implementation https://github.com/dDipankar/DeepPTM. Supplementary information Supplementary data are available at Bioinformatics online.

Baisya, Dipankar Ranjan (ORCID:0000000267847359)↗

Accuracy of predictions made by machine learned models for biocrude yields obtained from hydrothermal liquefaction of organic wastes

Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. In our report the models and analysis represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.

42 ENGINEERING↗

Genome‐wide association and genomic prediction for yield and component traits of Miscanthus sacchariflorus

Abstract Accelerating biomass improvement is a major goal of Miscanthus breeding. The development and implementation of genomic‐enabled breeding tools, like marker‐assisted selection (MAS) and genomic selection, has the potential to improve the efficiency of Miscanthus breeding. The present study conducted genome‐wide association (GWA) and genomic prediction of biomass yield and 14 yield‐components traits in Miscanthus sacchariflorus . We evaluated a diversity panel with 590 accessions of M. sacchariflorus grown across 4 years in one subtropical and three temperate locations and genotyped with 268,109 single‐nucleotide polymorphisms (SNPs). The GWA study identified a total of 835 significant SNPs and 674 candidate genes across all traits and locations. Of the significant SNPs identified, 280 were localized in mapped quantitative trait loci intervals and proximal to SNPs identified for similar traits in previously reported Miscanthus studies, providing additional support for the importance of these genomic regions for biomass yield. Our study gave insights into the genetic basis for yield‐component traits in M. sacchariflorus that may facilitate marker‐assisted breeding for biomass yield. Genomic prediction accuracy for the yield‐related traits ranged from 0.15 to 0.52 across all locations and genetic groups. Prediction accuracies within the six genetic groupings of M. sacchariflorus were limited due to low sample sizes. Nevertheless, the Korea/NE China/Russia ( N = 237) genetic group had the highest prediction accuracy of all genetic groups (ranging 0.26–0.71), suggesting that with adequate sample sizes, there is strong potential for genomic selection within the genetic groupings of M. sacchariflorus . This study indicated that MAS and genomic prediction will likely be beneficial for conducting population‐improvement of M. sacchariflorus .

09 BIOMASS FUELS↗

Data for Genome-Wide Association and Genomic Prediction for Yield and Component Traits of Miscanthus sacchariflorus

Accelerating biomass improvement is a major goal of miscanthus breeding. The development and implementation of genomic-enabled breeding tools, like marker-assisted selection (MAS) and genomic selection, has the potential to improve the efficiency of miscanthus breeding. The present study conducted genome-wide association (GWA) and genomic prediction of biomass yield and 14 yield-components traits in Miscanthus sacchariflorus . We evaluated a diversity panel with 590 accessions of M. sacchariflorus grown across four years in one subtropical and three temperate locations and genotyped with 268,109 single-nucleotide polymorphisms (SNPs). The GWA study identified a total of 835 significant SNPs and 674 candidate genes across all traits and locations. Of the significant SNPs identified, 280 were localized in mapped quantitative trait loci intervals and proximal to SNPs identified for similar traits in previously reported miscanthus studies, providing additional support for the importance of these genomic regions for biomass yield. Our study gave insights into the genetic basis for yield-component traits in M. sacchariflorus that may facilitate marker-assisted breeding for biomass yield. Genomic prediction accuracy for the yield-related traits ranged from 0.15 to 0.52 across all locations and genetic groups. Prediction accuracies within the six genetic groupings of M. sacchariflorus were limited due to low sample sizes. Nevertheless, the Korea/NE China/Russia (N = 237) genetic group had the highest prediction accuracy of all genetic groups (ranging 0.26–0.71), suggesting that with adequate sample sizes, there is strong potential for genomic selection within the genetic groupings of M. sacchariflorus . This study indicated that MAS and genomic prediction will likely be beneficial for conducting population-improvement of M. sacchariflorus .

Biomass Analytics↗

Predicting the formation of fractionally doped perovskite oxides by a function-confined machine learning method

Abstract Fractionally doped perovskites oxides (FDPOs) have demonstrated ubiquitous applications such as energy conversion, storage and harvesting, catalysis, sensor, superconductor, ferroelectric, piezoelectric, magnetic, and luminescence. Hence, an accurate, cost-effective, and easy-to-use methodology to discover new compositions is much needed. Here, we developed a function-confined machine learning methodology to discover new FDPOs with high prediction accuracy from limited experimental data. By focusing on a specific application, namely solar thermochemical hydrogen production, we collected 632 training data and defined 21 desirable features. Our gradient boosting classifier model achieved a high prediction accuracy of 95.4% and a high F1 score of 0.921. Furthermore, when verified on additional 36 experimental data from existing literature, the model showed a prediction accuracy of 94.4%. With the help of this machine learning approach, we identified and synthesized 11 new FDPO compositions, 7 of which are relevant for solar thermochemical hydrogen production. We believe this confined machine learning methodology can be used to discover, from limited data, FDPOs with other specific application purposes.

08 HYDROGEN↗

Di-CNN: Domain-Knowledge-Informed Convolutional Neural Network for Manufacturing Quality Prediction

In manufacturing, convolutional neural networks (CNNs) are widely used on image sensor data for data-driven process monitoring and quality prediction. However, as purely data-driven models, CNNs do not integrate physical measures or practical considerations into the model structure or training procedure. Consequently, CNNs’ prediction accuracy can be limited, and model outputs may be hard to interpret practically. This study aims to leverage manufacturing domain knowledge to improve the accuracy and interpretability of CNNs in quality prediction. A novel CNN model, named Di-CNN, was developed that learns from both design-stage information (such as working condition and operational mode) and real-time sensor data, and adaptively weighs these data sources during model training. It exploits domain knowledge to guide model training, thus improving prediction accuracy and model interpretability. A case study on resistance spot welding, a popular lightweight metal-joining process for automotive manufacturing, compared the performance of (1) a Di-CNN with adaptive weights (the proposed model), (2) a Di-CNN without adaptive weights, and (3) a conventional CNN. The quality prediction results were measured with the mean squared error (MSE) over sixfold cross-validation. Model (1) achieved a mean MSE of 6.8866 and a median MSE of 6.1916, Model (2) achieved 13.6171 and 13.1343, and Model (3) achieved 27.2935 and 25.6117, demonstrating the superior performance of the proposed model.

47 OTHER INSTRUMENTATION↗

Sub-millisecond keyhole pore detection in laser powder bed fusion using sound and light sensors and machine learning

Laser powder bed fusion is a mainstream additive manufacturing technology widely used to manufacture complex parts in prominent sectors, including aerospace, biomedical, and automotive industries. However, during the printing process, the presence of an unstable vapor depression can lead to a type of defect called keyhole porosity, which is detrimental to the part quality. In this study, we developed an effective approach to locally detect the generation of keyhole pores during the printing process by leveraging machine learning and a suite of optical and acoustic sensors. Simultaneous synchrotron x-ray imaging allows the direct visualization of pore generation events inside the sample, offering high-fidelity ground truth. A neural network model adopting SqueezeNet architecture using single-sensor data was developed to evaluate the fidelity of each sensor for capturing keyhole pore generation events. Our comparative study shows that the near infrared images gave the highest prediction accuracy, followed by 100 kHz and 20 kHz microphones, and the photodiode sensitive to processing laser wavelength had the lowest accuracy. Using a single sensor, over 90% prediction accuracy can be achieved with a temporal resolution as short as 0.1 ms. A data fusion scheme was also developed with features extracted using SqueezeNet neural network architecture and classification using different machine learning algorithms. Our work demonstrates the correlation between the characteristic optical and acoustic emissions and the keyhole oscillation behavior, and thereby provides strong physics support for the machine learning approach.

36 MATERIALS SCIENCE↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Bayesian Optimized Deep Ensemble for Uncertainty Quantification of Deep Neural Networks: a System Safety Case Study on Sodium Fast Reactor Thermal Stratification Modeling

Deep neural networks (DNNs) are increasingly important to scientific computing and engineering system simulations. Accurate uncertainty quantification (UQ) for DNNs is critical in safety-sensitive engineering domains. Traditional Deep Ensemble (DE) methods, while easy to implement, frequently suffer from poorly calibrated uncertainty estimates and limited predictive accuracy due to reliance on fixed architectures with varied weight initializations. To address these issues, we introduce a workflow that combines Bayesian Optimization (BO) and DE. The workflow is modular, scalable, and integrates parallel BO initialized with Sobol sequences to individually optimize the hyperparameters of each ensemble member. This method enhances ensemble diversity, improves predictive accuracy, and provides reliable uncertainty estimates. We evaluate the proposed BODE approach in a sodium fast reactor thermal stratification modeling case study, where we used a densely connected convolutional neural network to predict turbulent viscosity during the reactor transient with consideration of data noise. We benchmark its performance against several optimization approaches, including baseline deep ensemble, evolutionary algorithm-optimized ensemble, ensemble formed via random search combined with greedy selection, and a BO ensemble using random initialization. Here, our results demonstrate superior performance of the developed BODE approach. In noise-free scenarios, BODE notably reduces incorrect aleatoric uncertainty and significantly enhances predictive accuracy. Under conditions of 5% and 10% Gaussian noise, BODE adaptively quantifies uncertainty proportional to data noise, achieving up to an 80% reduction in root mean square error compared to baseline methods and producing well-calibrated prediction intervals.

Bayesian optimization↗

Towards a robust out-of-the-box neural network model for genomic data

The accurate prediction of biological features from genomic data is paramount for precision medicine and sustainable agriculture. For decades, neural network models have been widely popular in fields like computer vision, astrophysics and targeted marketing given their prediction accuracy and their robust performance under big data settings. Yet neural network models have not made a successful transition into the medical and biological world due to the ubiquitous characteristics of biological data such as modest sample sizes, sparsity, and extreme heterogeneity. Here, we investigate the robustness, generalization potential and prediction accuracy of widely used convolutional neural network and natural language processing models with a variety of heterogeneous genomic datasets. Mainly, recurrent neural network models outperform convolutional neural network models in terms of prediction accuracy, overfitting and transferability across the datasets under study. While the perspective of a robust out-of-the-box neural network model is out of reach, we identify certain model characteristics that translate well across datasets and could serve as a baseline model for translational researchers.

59 BASIC BIOLOGICAL SCIENCES↗