Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Evaluation of the EarthSHAB Stratospheric Solar Hot Air Balloon Flight Prediction Model Using Balloon Trajectory Data

Abstract The heliotrope is a solar balloon design which is constructed out of painter’s plastic, and the exterior is coated in charcoal powder. Darkening the plastic gives the balloon a high solar absorptance, which allows it to ascend into the lower stratosphere and float for hours at a time. The balloons have previously been used to lift scientific instruments into the stratosphere to study chemical explosions, earthquakes, and stratospheric aerosols. They have also been proposed as a platform for planetary exploration. Flight predictions are crucial to preflight planning to reduce safety risks and meet flight objectives. However, there exists a wide range of possible flight paths due to varying environmental conditions and solar balloon configurations. EarthSHAB is one such software that was designed to support flight planning using the weather forecasts and balloon properties to predict the flight path of a solar balloon. We compare EarthSHAB-simulated flight paths to a set of observed flight paths for the 3.5-m diameter heliotrope design called the “Cloudskimmer.” Using the criteria that the modeled paths must fall within 5% of the observations to be considered successful, we found that EarthSHAB successfully predicted the Cloudskimmer ascent rate and average float altitude 10% and 90% of the time, respectively. We also found that the average difference in the observed and predicted landing locations was 97 km and landing times were 54 ± 38 min. Significant deviations between the observed and predicted ascent rates and excursions at float were found to be associated with heavy payloads and convective cloud development, respectively. Significance Statement Solar balloons are used to lift scientific instruments into the lower stratosphere for hours at a time to study chemical explosions, earthquakes, stratospheric aerosols, and more. The flight paths of solar balloons can be difficult to predict due to variability in their design and surrounding environment. We evaluate the accuracy of EarthSHAB, a software that predicts the altitude profile and horizontal trajectory of a solar balloon using inputs such as the weather, balloon size, and balloon mass. Our results suggest that for the balloon design used in this study, EarthSHAB is best suited for modeling the behavior of balloons with lightweight payloads that do not fly within or directly above clouds.

Lien, Jessica M. [Sandia National Laboratories, Al↗

Subseasonal prediction with and without a well-represented stratosphere in CESM1

There is a growing demand for understanding sources of predictability on subseasonal to seasonal (S2S) time scales. Predictability at subseasonal time scales is believed to come from processes varying slower than the atmosphere such as soil moisture, snowpack, sea ice, and ocean heat content. The stratosphere as well as tropospheric modes of variability can also provide predictability at subseasonal time scales. Furthermore, the contributions of the above sources to S2S predictability are not well quantified. Here we evaluate the subseasonal prediction skill of the Community Earth System Model, version 1 (CESM1), in the default version of the model as well as a version with the improved representation of stratospheric variability to assess the role of an improved stratosphere on prediction skill. We demonstrate that the subseasonal skill of CESM1 for surface temperature and precipitation is comparable to that of operational models. We find that a better-resolved stratosphere improves stratospheric but not surface prediction skill for weeks 3–4.

54 ENVIRONMENTAL SCIENCES↗

Evaluating Ensemble Predictions of South Asian Monsoon Low Pressure System Genesis

Abstract Synoptic-scale vortices known as monsoon low pressure systems (LPSs) frequently produce intense precipitation and hydrological disasters in South Asia, so accurately forecasting LPS genesis is crucial for improving disaster preparedness and response. However, the accuracy of LPS genesis forecasts by numerical weather prediction models has remained unknown. Here, we evaluate the performance of two global ensemble models—the U.S. Global Ensemble Forecast System (GEFS) and the Ensemble Prediction System of the European Centre for Medium-Range Weather Forecasts (ECMWF)—in predicting LPS genesis during the years 2021–22. The GEFS successfully predicted about half the observed LPS genesis events 1–2 days in advance; the ECMWF model captured an additional 10% of observed genesis events. Both models had a false alarm ratio (FAR) of around 50% for 1–2-day lead times. In both ensembles, the control run typically exhibited a higher probability of detection (POD) of observed events and a lower FAR compared to the perturbed ensemble members. However, a consensus forecast, in which genesis is predicted when at least 20% of ensemble members forecast LPS formation, had POD values surpassing those of the control run for all lead times. Moreover, probabilistic predictions of genesis over the Bay of Bengal, where most LPSs form, were skillful, with the fraction of ensemble members predicting LPS formation over a 5-day lead time approximating the observed frequency of genesis, without any adjustment or bias correction.

Suhas, D. L.↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

Predicting potential adverse events using safety data from marketed drugs

Abstract Background While clinical trials are considered the gold standard for detecting adverse events, often these trials are not sufficiently powered to detect difficult to observe adverse events. We developed a preliminary approach to predict 135 adverse events using post-market safety data from marketed drugs. Adverse event information available from FDA product labels and scientific literature for drugs that have the same activity at one or more of the same targets, structural and target similarities, and the duration of post market experience were used as features for a classifier algorithm. The proposed method was studied using 54 drugs and a probabilistic approach of performance evaluation using bootstrapping with 10,000 iterations. Results Out of 135 adverse events, 53 had high probability of having high positive predictive value. Cross validation showed that 32% of the model-predicted safety label changes occurred within four to nine years of approval (median: six years). Conclusions This approach predicts 53 serious adverse events with high positive predictive values where well-characterized target-event relationships exist. Adverse events with well-defined target-event associations were better predicted compared to adverse events that may be idiosyncratic or related to secondary target effects that were poorly captured. Further enhancement of this model with additional features, such as target prediction and drug binding data, may increase accuracy.

Daluwatte, Chathuri↗

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Biochemistry & Molecular Biology↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Predicting the viability of beta-lactamase: How folding and binding free energies correlate with beta-lactamase fitness

One of the long-standing holy grails of molecular evolution has been the ability to predict an organism’s fitness directly from its genotype. With such predictive abilities in hand, researchers would be able to more accurately forecast how organisms will evolve and how proteins with novel functions could be engineered, leading to revolutionary advances in medicine and biotechnology. In this work, we assemble the largest reported set of experimental TEM-1 β-lactamase folding free energies and use this data in conjunction with previously acquired fitness data and computational free energy predictions to determine how much of the fitness of β-lactamase can be directly predicted by thermodynamic folding and binding free energies. We focus upon β-lactamase because of its long history as a model enzyme and its central role in antibiotic resistance. Based upon a set of 21 β-lactamase single and double mutants expressly designed to influence protein folding, we first demonstrate that modeling software designed to compute folding free energies such as FoldX and PyRosetta can meaningfully, although not perfectly, predict the experimental folding free energies of single mutants. Interestingly, while these techniques also yield sensible double mutant free energies, we show that they do so for the wrong physical reasons. We then go on to assess how well both experimental and computational folding free energies explain single mutant fitness. We find that folding free energies account for, at most, 24% of the variance in β-lactamase fitness values according to linear models and, somewhat surprisingly, complementing folding free energies with computationally-predicted binding free energies of residues near the active site only increases the folding-only figure by a few percent. This strongly suggests that the majority of β-lactamase’s fitness is controlled by factors other than free energies. Overall, our results shed a bright light on to what extent the community is justified in using thermodynamic measures to infer protein fitness as well as how applicable modern computational techniques for predicting free energies will be to the large data sets of multiply-mutated proteins forthcoming.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improving five-year survival prediction via multitask learning across HPV-related cancers

Oncology is a highly siloed field of research in which sub-disciplinary specialization has limited the amount of information shared between researchers of distinct cancer types. This can be attributed to legitimate differences in the physiology and carcinogenesis of cancers affecting distinct anatomical sites. However, underlying processes that are shared across seemingly disparate cancers probably affect prognosis. The objective of the current study is to investigate whether multitask learning improves 5-year survival cancer patient survival prediction by leveraging information across anatomically distinct HPV related cancers. Furthermore, data were obtained from the Surveillance, Epidemiology, and End Results (SEER) program database. The study cohort consisted of 29,768 primary cancer cases diagnosed in the United States between 2004 and 2015. Ten different cancer diagnoses were selected, all with a known association with HPV risk. In the analysis, the cancer diagnoses were categorized into three distinct topography groups of varying specificity. The most specific topography grouping consisted of 10 original cancer diagnoses differentiated by the first two digits of the ICD-O-3 topography code. The second topography grouping consisted of cancer diagnoses categorized into six distinct organ groups. Finally, the third topography grouping consisted of just two groups, head-neck cancers and ano-genital cancers. The tasks were to predict 5-year survival for patients within the different topography groups using 14 predictive features which were selected among descriptive variables available in the SEER database. The information from the predictive features was shared between tasks in three different ways, resulting in three distinct predictive models: 1) Information was not shared between patients assigned to different tasks (single task learning); 2) Information was shared between all patients, regardless of task (pooled model); 3) Only relevant information was shared between patients grouped to different tasks (multitask learning). Prediction performance was evaluated with Brier scores. All three models were evaluated against one another on each of the three distinct topography-defined tasks. The results showed that multitask classifiers achieved relative improvement for the majority of the scenarios studied compared to single task learning and pooled baseline methods. In this study, we have demonstrated that sharing information among anatomically distinct cancer types can lead to improved predictive survival models.

59 BASIC BIOLOGICAL SCIENCES↗

Unstructured clinical notes within the 24 hours since admission predict short, mid & long-term mortality in adult ICU patients

Mortality prediction for intensive care unit (ICU) patients is crucial for improving outcomes and efficient utilization of resources. Accessibility of electronic health records (EHR) has enabled data-driven predictive modeling using machine learning. However, very few studies rely solely on unstructured clinical notes from the EHR for mortality prediction. In this work, we propose a framework to predict short, mid, and long-term mortality in adult ICU patients using unstructured clinical notes from the MIMIC III database, natural language processing (NLP), and machine learning (ML) models. Depending on the statistical description of the patients’ length of stay, we define the short-term as 48-hour and 4-day period, the mid-term as 7-day and 10-day period, and the long-term as 15-day and 30-day period after admission. We found that by only using clinical notes within the 24 hours of admission, our framework can achieve a high area under the receiver operating characteristics (AU-ROC) score for short, mid and long-term mortality prediction tasks. The test AU-ROC scores are 0.87, 0.83, 0.83, 0.82, 0.82, and 0.82 for 48-hour, 4-day, 7-day, 10-day, 15-day, and 30-day period mortality prediction, respectively. We also provide a comparative study among three types of feature extraction techniques from NLP: frequency-based technique, fixed embedding-based technique, and dynamic embedding-based technique. Lastly, we provide an interpretation of the NLP-based predictive models using feature-importance scores.

60 APPLIED LIFE SCIENCES↗

Ising-Traffic: Using Ising Machine Learning to Predict Traffic Congestion under Uncertainty

This paper addresses the challenges in accurate and realtime traffic congestion prediction with uncertainty by proposing Ising-Traffic, a novel quantum-inspired dual-model Ising based traffic prediction framework which delivers higher accuracy and lower latency than SOTA solutions. While traditional and deep learning methods face the trade-off between algorithm complexity and computational efficiency, our Ising-based method leverages Ising’s inherent and unique capability of finding the state of a system with the lowest energy and applying it to traffic prediction. In this work, traffic prediction under uncertainty is formulated into two separate Ising models: Reconstruct-Ising and Predict-Ising. Reconstruct-Ising is mapped onto modern Ising machine and handles uncertainty in traffic accurately with negligible latency and energy consumption, while Predict-Ising is mapped onto traditional processors and predicts future congestion precisely with only at most 1.8% computational demands of existing solutions. Our evaluation shows Ising-Traffic delivers on average 98× speedups and 5% accuracy improvement over SOTA.

traffic flow control, Ising↗

A Spatiotemporal-Aware Weighting Scheme for Improving Climate Model Ensemble Predictions

Multimodel ensembling has been widely used to improve climate model predictions, and the improvement strongly depends on the ensembling scheme. In this work, we propose a Bayesian neural network (BNN) ensembling method, which combines climate models within a Bayesian model averaging framework, to improve the predictive capability of model ensembles. Our proposed BNN approach calculates spatiotemporally varying model weights and biases by leveraging individual models' simulation skill, calibrates the ensemble prediction against observations by considering observation data uncertainty, and quantifies epistemic uncertainty when extrapolating to new conditions. More importantly, the BNN method provides interpretability about which climate model contributes more to the ensemble prediction at which locations and times. Thus, beyond its predictive capability, the method also brings insights and understanding of the models to guide further model and data development. In this study, we design experiments using an ensemble of CMIP6 climate model simulations to illustrate the BNN ensembling method's capability with respect to prediction accuracy, interpretability, and uncertainty quantification (UQ). We demonstrate that BNN can correctly assign larger weights to the regions and seasons where the individual model fits the observation better. Moreover, its offered interpretability is consistent with our understanding of localized climate model performance. Additionally, BNN shows an increasing uncertainty when the prediction is farther away from the period with constrained data, which appropriately reflects our trustworthiness of the models in the changing climate.

54 ENVIRONMENTAL SCIENCES↗

Force Balance Model Assessment for Mechanistic Prediction of Sliding Bubble Velocity in Vertical Subcooled Boiling Flow

The bubble sliding after departing from nucleation site is frequently observed in flow boiling systems and, the crucial impact on wall heat transfer has been evidenced through many experiments. As a result, the heat transfer modeling associated with sliding bubble has become one of the subjects of great attention in CFD boiling heat transfer community. The modeling efforts are primarily aimed at improving the existing Heat Flux Partitioning (HFP) model via the implementation of sliding bubble-induced heat transfer. The performance of HFP model depends inherently on the fidelity of sub-models used to predict the fundamental bubble parameters (e.g., bubble departure/lift-off diameter). In the same context, the accurate prediction of sliding bubble parameters (e.g., sliding bubble growth, sliding bubble velocity) is essential to achieving the successful heat transfer modeling associated with sliding bubble. Of many sliding bubble parameters, this paper deals with the sliding bubble velocity. Specifically, the force balance model was assessed in view of the sliding bubble velocity prediction. The parametric effect of key sub-models (e.g., drag force, bubble growth models) used in the force balance equation was investigated. The experimental data from Maity (2000) and Yoo et al. (2016) were used for demonstrating the model predictive performance. It was found that the force balance model proposed in this study was able to predict well the bubble sliding velocity based on the accurate prediction of bubble growth during sliding. The predictive performance was proven with the experimental data measured under various subcooled boiling conditions of both water and refrigerant (NOVEC-7000) flowing upward in vertical channels.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Robust Molecular Predictive Methods for Novel Polymer Discovery and Applications

Polymeric materials are ubiquitous in modern society and they play an instrumental role in almost all industries, undoubtedly including the energy and environment sectors. Increased demand of energy and awareness to sustainability both necessitates the development of novel polymers with enhanced properties. Unfortunately, their structural and behavioral complexity render such discovery challenging and impeded. To address this problem, scientists are developing various computational modeling techniques and leveraging their power to depict the relationship between structural characteristics of polymers and their properties (such as rheological behaviors), and use such prediction to guide the design and syntheses of novel polymeric materials with enhanced performances. Unfortunately, predicting the relationships between polymer structure and composition with rheological properties via atomistic modeling is still a major challenge because of the extended time and length scales involved. Studying dynamic shear viscosity and linear viscoelasticity using molecular models requires capabilities that have been elusive, including representation of large molecular weight chains with an effective internal scale capable of describing entanglement, shear-rates that are in the s-1 scale with accurate quantitative stresses, and chemically-realistic combinations of both homogeneous and heterogeneous systems. Motivated by these unmet challenges, the overall technical objective of this DOE-STTR Phase II project is to develop robust molecular predictive methods for advanced polymer discovery and applications and especially for designing and demonstrating the “smart” polymer-based waterflooding enhanced oil recovery (EOR) process. In particular, we apply state-of-the-art molecular modeling methods developed by our academic partner, Materials Stimulation Center (MSC) at California Institute of Technology (Caltech), to facilitate and accelerate the experimental discovery processes. During the Phase I of this project, we had focused on development and demonstration of the molecular modeling methods to describe rheological properties of non-Newtonian polymer fluids, and to improve our fundamental understandings of shear-thickening mechanism and kinetics. In Phase II, we further apply the theoretical models to guide our experimental programs to improve our design of smart rheology modifier (SRM) polymers and their optimization for EOR. Specifically, we have three objectives in the Phase II study: (1) to further improve out computational modeling methods, coupling with the advanced machine learning algorithms; (2) to develop cost-effective and efficient SRM-flooding process suitable for EOR applications under typical reservoir conditions; and (3) to further explore the application of our molecular predictive models for innovative material discovery in other industrial applications. The recent development of our multiscale predictive framework allows the successful prediction of rheological properties from the chemical structure for polymers of experimentally relevant molecular weights, and provides an in-silico machine learning engine for screening novel compositions and structures with optimized non-Newtonian response, required for both shear-thinning and shear-thickening applications. Our framework provides: (1) procedures and tools for systematic coarsening from atomistic models and reverse mapping of coarse-grain models to atomistic, (2) unique ab initio methods to characterize the atomistic origin of colloidal and interfacial interactions and phenomena, (3) systematic structure and composition builders based on practical descriptors that drive rheological changes in polymer melts and diluted polymer mixtures, (4) a rheological properties engine capable of predicting viscosity in the zero-shear limit and under realistic dynamic conditions (for shear-rates commensurate with experiments) for large heterogeneous systems, (5) coarse-grain force fields with improved non-bond descriptions based on accurate quantum mechanics, (6) an in-silico screening machine learning engine that feeds from the systematic model builders to cover the descriptors search space, computes the rheological properties from converged trajectories spanning sub-milliseconds and ranks them for each structure/composition using an automated viscosity-vs-shear rate fitness function that can be tuned for shear-thickening, shear-thinning and other rheological responses.

02 PETROLEUM↗

Probabilistic Predictions for Fastener Failure in the Sandia Mechanics Challenge Using the Discrete-Direct Uncertainty Quantification Approach

This paper documents the blind and post-blind analysis predictions for the 2023 Sandia Mechanics Challenge (SMC), which involved predicting the behavior of a threaded fastener joint structure subjected to shock loading. Utilizing repeat sets of fastener calibration data from various experimental configurations including tension, double shear, and joint tension, we developed a library of calibrated models which were propagated through the application model using the Discrete-Direct (DD) uncertainty quantification (UQ) approach. Although the initial blind predictions did not incorporate spare-sample processing to quantify fastener failure probabilities, the analyses yielded reasonable conclusions aligned with experimental results. In the post-blind analysis phase, we focused on enhancing the fidelity of the aluminum constitutive model and innovating the DD approach to obtain probabilistic predictions for fastener failure, particularly when quantities of interest (QoIs) approach their bounds. The improved aluminum model captures the behavior of the cantilever under shock loading more accurately, predicting both partial and complete cracks, although it tends to underpredict failure propagation. The enhanced DD approach facilitates probabilistic predictions that reflect the interdependent failure mechanisms of the fasteners and the cantilever, revealing that while certain fasteners are more likely to fail, the failure does not necessarily follow a progressive pattern. Overall, the post-blind analyses significantly improved the predictive capabilities of the model, providing valuable insights into the SMC application and establishing a robust foundation for informed engineering decisions. The methodology demonstrates a cost-effective and extensible approach suitable for a wide range of applications, highlighting the importance of uncertainty quantification to provide context for engineering decision making.

42 ENGINEERING↗