Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The Solar Influencer Next Door: Predicting Low-Income Solar Referrals and Leads

Increasing the adoption of solar among low-to-moderate income (LMI) households remains an important policy goal because of its promise to simultaneously reduce energy burden and support the just distribution of benefits of renewable energy. However, scaling LMI solar remains challenging due to affordability and access issues. Most existing LMI adoption has occurred under public-funded programs, highlighting the importance of increasing the cost-effectiveness of these programs at scale. We develop a new household-level data set on LMI solar lead acquisition, referrals, and adoption to understand the processes through which LMI solar uptake has occurred in California. Then, we develop models to predict two sub-mechanisms in the solar adoption process: whether an otherwise qualified lead becomes "lost" i.e. non-responsive to outreach and, for existing clients, whether they refer solar to others. For the program analyzed, participants received their solar system at no cost, which deemphasizes economic drivers of solar adoption and could differ from other program experiences. Both models substantially improved the accuracy of prediction relative to a baseline. Overall, we find that peer effects and solar economics are important to predicting referrals, and household demographic factors in lead loss prediction. Finally, we find that referrals are both the highest quality and largest source of LMI solar leads, providing a promising mechanism to expand LMI programs further.

customer acquisition costs↗

A review of machine learning in building load prediction

The surge of machine learning in recent years has been empowering engineer modeling in various fields. The decreasing hardware cost, increasing data accessibility, and advances of building automation system (BAS) allow the collection and storage of a significant amount of building operation data. The two facts provide great opportunities of applying machine learning to building energy systems modeling and analysis. There are a great number of research papers on this topic but there lacks a comprehensive and general review to summarize the current development, limitations, gaps and future trend. In this review paper series, machine learning techniques in building energy system modeling and analysis are reviewed under the organization and logic of the machine learning definition by Tom M. Mitchell: a computer program is said to learn from experience E with respect to some class of tasks T and performance measure P if its performance at tasks in T, as measured by P, improves with experience E. This paper is the first part of the review paper series, which focuses on building load prediction. First, the applications of building load prediction model (task T) are reviewed. Then, the modeling algorithms improving machine learning performance and accuracy (performance P) are reviewed. At the same time, the literature on the data perspective for modeling (experience E), including data engineering from sensors level to data level, pre-processing, feature extraction and selection, is reviewed. Finally, what is well-studied and what is lacking but with great potential are concluded; the gaps between present and future utilization of machine learning techniques are identified; the future trend and development are also predicted. The target readers of this paper are not only researchers from the building side who can get exposed to cutting edge machine learning tools, but also those from machine learning side who can understand the potential and challenge to apply machine learning in buildings.

Liang, Zhang↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

A machine learning approach for clinker quality prediction and nonlinear model predictive control design for a rotary cement kiln

Abstract Cement manufacturing is energy‐intensive (5Gj/t) and comprises a significant portion of the energy footprint of concrete systems. Incorporating modern monitoring, simulation and control systems will allow lower energy use, lower environmental impact, and lower costs of this widely used construction material. One of the goals of the CESMII roadmap project on the Smart Manufacturing of Cement included developing an analytical process model for clinker quality that includes the chemistry of the kiln feed and accounts for critical process variables. This predictive model will be used in nonlinear model predictive control system designed to significantly reduce process energy use while maintaining or improving product quality. In the cement manufacturing plant used in this study, the kiln feed (meal) is tested every 12 h and used to estimate the mineral composition of the cement kiln output (clinker) using the stoichiometry‐based Bogue's model and the expertise of the plant operators. During kiln operation, kiln output (clinker) is sampled and tested every 2 h to measure its chemical and mineral composition. The predicted and measured values of the clinker composition are used by the plant operators to adjust the kiln input stream and the production process characteristics to maintain stable operation and uniform product quality. However, the time delay between prediction and testing, along with inaccuracies inherent in the Bogue's model have made any process changes designed to minimize energy use problematic, especially in‐light of potential clinker quality issues that process changes often pose. A new analytical model that integrates quality information and process operation information has been developed from data collected from 2 years of production from an operating cement facility. To make the model fuel‐type‐independent, consumed heat energy was computed in the model instead of fuel type and amount. A Feedforward Network was trained and tailored from collected data. Many data‐based simulations were conducted to quantitatively evaluate the proposed model and the 5‐fold cross‐validation procedure was used to test the models. The resulting predictive model was shown to have a low root mean square error (MSE) with respect to the estimated clinker mineral composition compared to that using the industry standard “Bogue’ model”. The end goal of this work was to develop a single machine learning tool that allows the use of quality control data and process control variables to improve energy efficiency of the process in a continuous fashion. The proposed nonlinear model predictive control system (NMPC) can generate predicted kiln production characteristics based on manipulated variables in manner that accurately follows the target product quality values. Simulation results also show that the proposed model produced accurate predictions of kiln outputs that fell within the required constraints, while manipulating control variables within typical operational ranges.

Ali, Asem M.↗

Prediction and Predictability of the Madden Julian Oscillation in the NASA GEOS-5 Seasonal-to-Subseasonal System

In this study, we examine the prediction skill and predictability of the Madden Julian Oscillation (MJO) in a recent version of the NASA GEOS-5 atmosphere-ocean coupled model run at at 1/2 degree horizontal resolution. The results are based on a suite of hindcasts produced as part of the NOAA SubX project, consisting of seven ensemble members initialized every 5 days for the period 1999-2015. The atmospheric initial conditions were taken from the Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2), and the ocean and the sea ice were taken from a GMAO ocean analysis. The land states were initialized from the MERRA-2 land output, which is based on observation-corrected precipitation fields. We investigated the MJO prediction skill in terms of the bivariate correlation coefficient for the real-time multivariate MJO (RMM) indices. The correlation coefficient stays at or above 0.5 out to forecast lead times of 26-36 days, with a pronounced increase in skill for forecasts initialized from phase 3, when the MJO convective anomaly is located in the central tropical Indian Ocean. A corresponding estimate of the upper limit of the predictability is calculated by considering a single ensemble member as the truth and verifying the ensemble mean of the remaining members against that. The predictability estimates fall between 35-37 days (taken as forecast lead when the correlation reaches 0.5) and are rather insensitive to the initial MJO phase. The model shows slightly higher skill when the initial conditions contain strong MJO events compared to weak events, although the difference in skill is evident only from lead 1 to 20. Similar to other models, the RMM-index-based skill arises mostly from the circulation components of the index. The skill of the convective component of the index drops to 0.5 by day 20 as opposed to day 30 for circulation fields. The propagation of the MJO anomalies over the Maritime Continent does not appear problematic in the GEOS-5 hindcasts implying that the Maritime Continent predictability barrier may not be a major concern in this model. Finally, the MJO prediction skill in this version of GEOS-5 is superior to that of the current seasonal prediction system at the GMAO; this could be partly attributed to a slightly better representation of the MJO in the free running version of this model and partly to the improved atmospheric initialization from MERRA-2.

Achuthavarier, Deepthi↗

A Machine Learning Approach to Predict Aircraft Landing Times using Mediated Predictions from Existing Systems

We developed a novel approach for predicting the landing time of airborne flights in real-time operations. The first step predicts a landing time by using mediation rules to select from among physics-based predictions (relying on the expected flight trajectory) already available in real time in the Federal Aviation Administration System Wide Information Management system data feeds. The second step uses a machine learning model built upon the mediated predictions. The model is trained to predict the error in the mediated prediction, using features describing the current state of an airborne flight. These features are calculated in real time from a relatively small number of data elements that are readily available for airborne flights. Initial results based on five months of data at six large airports demonstrate that incorporating a machine learning model on top of the mediated physics-based prediction can lead to substantial additional improvements in prediction quality.

Machine learning↗

A Machine Learning Approach to Predict Aircraft Landing Times using Mediated Predictions from Existing Systems

We developed a novel approach for predicting the landing time of airborne flights in real-time operations. The first step predicts a landing time by using mediation rules to select from among physics-based predictions (relying on the expected flight trajectory) already available in real time in the Federal Aviation Administration System Wide Information Management system data feeds. The second step uses a machine learning model built upon the mediated predictions. The model is trained to predict the error in the mediated prediction, using features describing the current state of an airborne flight. These features are calculated in real time from a relatively small number of data elements that are readily available for airborne flights. Initial results based on five months of data at six large airports demonstrate that incorporating a machine learning model on top of the mediated physics-based prediction can lead to substantial additional improvements in prediction quality.

Machine learning↗

Comparative Study on the Machine Learning-Based Prediction of Adsorption Energies for Ring and Chain Species on Metal Catalyst Surfaces

Computation of adsorption and transition state energies for a large number of surface intermediates for numerous active site models pose significant computational overhead in computational screening of catalysts. Machine learning (ML) techniques can be used to predict part of these energies. To predict the energies, ML models need to be fed appropriate metal and species descriptors. For complex surface chemistries, the structures of the intermediate species can vary greatly. In this paper, working with the hydrodeoxygenation of succinic acid on six different metal surfaces, we have studied the effect of linear and non-linear ML models used along with pen-and-paper based species descriptors and two categories of metal descriptors on two different categories of intermediate species: chain and ring. More specifically, our computations include the prediction of chain species when trained on only chain species and also when trained on both chain and ring species. Similar computations were performed for predictions of ring species. In each case, results of linear ML models were compared with kernel based non-linear models. Our results indicate that ring species data does not improve the prediction of chain species. Similarly, chain species data does not improve the prediction of ring species. The use of non-linear ML models, however, did help to minimize the prediction errors compared to the linear models. Furthermore, the study also shows that electronic or adsorption energy based metal descriptors along with bond count based species fingerprints can achieve a mean absolute error (MAE) of less than 0.2 eV for complex chain molecules when used with an appropriate machine learning model.

Adsorption↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Observational Needs for Improving Ocean and Coupled Reanalysis, S2S Prediction, and Decadal Prediction

Developments in ocean data assimilation (DA) and observing system technologies are intertwined. New observation types lead to new DA methods, and new DA methods such as Coupled Data Assimilation can change the value of existing observations or indicate where new observations can have greater utility for monitoring and prediction. Practitioners are encouraged to make better use of observations that are already available, for example in strongly coupled data assimilation where ocean observations can be used to improve atmospheric analyses and vice versa. Ocean reanalyses are useful for the analysis of climate,as well as initializing operational long-range prediction models. There are remaining challenges for ocean reanalyses due to biases and abrupt changes in the ocean observing system throughout its history, the presence of biases and drifts in models, and simplifying assumptions made in the DA methods. From a governance point of view, more support is needed to interface the observing community and the ocean DA community. For prediction applications, the ocean DA community must work with the ocean observing community to establish protocols for rapid communication of ocean observing data on NWP timescales. There is potential for new observations to enhance the observing system by supporting prediction on multiple timescales, ranging from the typical timescale of numerical weather prediction covering hours to weeks, out to multiple decades. It is highly encouraged that communication be fostered between thesecommunities to allow operational prediction centers the ability to provide guidance to the design of a sustained and adaptive observing network.

Ocean reanalysis↗

FAST.Farm Development and Validation of Structural Load Prediction Against Large Eddy Simulations

FAST.Farm is a mid-fidelity engineering tool developed by the National Renewable Energy Laboratory targeted at accurately and efficiently predicting wind turbine power production and structural loading in wind farm settings, including wake interactions between turbines. FAST.Farm is based on several principles of the dynamic wake meandering (DWM) model, but also addresses limitations of previous DWM implementations. Previous FAST.Farm studies have shown the similarities and differences between FAST.Farm and large eddy simulations for rigid turbine cases. The objective of this work is to quantify the ability of FAST.Farm to accurately predict turbine structural response in a small wind farm. This is done by comparing FAST.Farm structural response to SOWFA-OpenFAST results for three laterally-aligned turbines. The purpose of this study is to characterize the similarities and differences between FAST.Farm and a higher fidelity model for predicting turbine structural response in a wind farm for differing atmospheric inflows. Strong statistical agreement was found between FAST.Farm and SOWFA-OpenFAST structural response for the non-waked upstream turbine, and good agreement was found for the downstream turbines for most structural quantities. Higher differences were seen for downstream turbines with low ambient turbulence intensity or yawed turbines, suggesting areas for FAST.Farm wake dynamics modeling improvements. For all cases and turbines, small statistical differences were seen between blade deflections and bending moments, with larger differences for tower-top and tower-base bending moments. Overall, the results establish confidence for applying FAST.Farm to wind farm power and loads analyses and identify areas where further model validations and model improvements should be targeted.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Thousands of small, novel genes predicted in global phage genomes

Small genes (<150nucleotides) have been systematically overlooked in phage genomes. We employ a large scale comparative genomics approach to predict >40,000 small-gene families in 2.3 million phage genome contigs. We find that small genes in phage genomes are approximately 3-fold more prevalent than in host prokaryotic genomes. Our approach enriches for small genes that are translated in microbiomes, suggesting the small genes identified are coding. More than 9,000 families encode potentially secreted or transmembrane proteins, more than 5,000families encode predicted anti-CRISPR proteins, and more than500families encode predicted antimicrobial proteins. By combining homology and genomic-neighborhood analyses, we reveal substantial novelty and diversity within phage biology, including small phage genes found in multiple host phyla, small genes encoding proteins that play essential roles in host infection, and small genes that share genomic neighborhoods and whose encoded proteins may share related functions.

Fremin, Brayon↗

The Cell Utilized Partitioning Model as a Predictive Tool for Optimizing Counter-Current Chromatography Processes

Counter-current chromatography (CCC) is capable of unique elution modes that isolate analytes using the movement of the stationary phase in addition to moving the mobile phase. These modes include elution-extrusion CCC (EECCC) and dual-mode CCC (DM CCC) that are not possible in traditional solid-liquid chromatography systems. Although EECCC and DM CCC are widely used to recover highly retained components, to our knowledge, optimizing the elution process in these modes with predictive models has not been reported. To address this gap, we developed a predictive model for CCC dubbed the Cell Utilized Partitioning (CUP) model. The CUP model accurately predicts the effluents of multicomponent separations in EECCC and DM CCC modes when compared to experimental data. Furthermore, CUP model simulations were extended to investigate the influence of operating and intrinsic parameters on the yield and productivity, and to compare the separation performances of EECCC and DM CCC in various conditions. The results demonstrate that low distribution constants, usually a KD less than 1, and a selectivity > 1.3, under specific flowrate ranges, increase both productivity and yield. From these results, generalized optimization and scaleup guidelines are proposed that can apply to research settings and to industrial processes to maximize preparative CCC performance.

BIOMASS FUELS,INORGANIC, ORGANIC, PHYSICAL, AND AN↗

Prediction of Redox Potentials for U, Np, Pu, and Am in Aqueous Solution

The redox properties of the actinides in aqueous solution are important for fuel production/reprocessing and understanding the environmental impact of nuclear waste. The redox potentials for U, Np, Pu, and Am in oxidation states from 0 up to VII (as appropriate) in aqueous solutions have been predicted at the density functional theory level with the B3LYP functional, Stuttgart small core pseudopotential basis sets for the actinides, and explicit (30H 2 O molecules)/implicit treatment of the aqueous solvent using the self-consistent reaction field COSMO and SMD approaches for the implicit solvation. The predictions of the structural parameters of clusters incorporating first and second solvation shells are consistent with the available experimental data., Our results are typically within 0.2 V of the available experimental data using two explicit solvation shells with an implicit solvent model. The use of the PW91 functional substantially improved the prediction of the Pu(VI/V) redox couple. The redox couples for An(VI/IV) and An(V/IV) which involve the addition of protons and removal of the actinyl oxygens led to slightly larger differences from experiment. Here, the An(IV/0) and An(III/0) couples were reliably predicted with our approach. Predictions of the unknown An(II/I) redox potentials were negative, consistent with expectations, and predictions for unknown An(VII/VI), An(III/II), and An(II/0) redox couples improve prior estimates.

Actinides↗

Short-Term Load Forecasting Considering EV Charging Loads with Prediction Interval Evaluation

Short-term load forecasting plays a critical role in power system planning and operation. Along with the electrification of various loads, electricity demands are becoming increasingly hard to predict. Notably, the recent rise in electric vehicles (EVs) has further contributed to this unpredictability. To address this issue, this paper proposes a probabilistic load forecasting strategy utilizing Gaussian process regression, structured in a day-ahead manner. While many works focus on deterministic prediction, probabilistic forecasting offers additional insights into variability and uncertainty, enabling more flexible and reliable operation for power systems. To enhance the accuracy of the load forecasting model, the inputs include features related to EV charging habits as well as commonly used weather information. The load forecasting results are evaluated using various metrics, including conventional ones that assess the accuracy of point forecasts, as well as additional metrics that test the reliability of prediction intervals. The proposed load forecasting method is finally tested on real residential power consumption data and EV charging data sampled from real-world sources. The results prove that the new features can greatly improve the performance of the load forecasting method.

electrical vehicle↗

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

Energy Prediction Impact of the Space Level Occupancy Schedule for a Primary School

By using the same occupant schedule for all spaces, a building level occupancy schedule in building energy modelling can reduce the cost and time of data collection, especially for large-scale simulations or when detailed occupancy data cannot be obtained. However, by describing the unique occupancy status in each space, a space level schedule can better reflect real-world scenarios. This research investigates the energy prediction impact of a space level occupancy schedule for primary school modelling in 16 ASHRAE climate zones. The results show that when switched from a building level to a space level occupancy schedule, the energy prediction difference is between -1.0% and 0.8%. Generally, the predicted building energy consumption using a space level occupancy schedule is higher than when using a building level occupancy schedule in hot and warm climate zones, but lower in other climate zones.

building energy model↗