Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Quantitative assessment of environmental phenomena on maximum pit size predictions in marine environments

Maximum pit sizes were predicted for dilute and concentrated NaCl and MgCl 2 solutions as well as sea-salt brine solutions corresponding to 40% relative humidity (RH) (MgCl 2 -rich) and 76% RH (NaCl-rich) at 25 °C. A quantitative method was developed to capture the effects of various cathode evolution phenomena including precipitation and dehydration reactions. Additionally, the sensitivity of the model to input parameters was explored. Despite one's intuition, the highest chloride concentration (roughly 10.3 M Cl – ) did not produce the largest predicted pit size as the ohmic drop was more severe in concentrated MgCl 2 solutions. Therefore, the largest predicted pits were calculated for saturated NaCl (roughly 5 M Cl – ). Next, it was determined that pit size predictions are most sensitive to model input parameters for concentrated brines. However, when the effects of cathodic reactions on brine chemistry are considered, the sensitivity to input parameters is decreased. Although there was not one main input parameter that influenced pit size predictions, two main categories were identified. Under similar chloride concentrations (similar RH), the water layer thickness (WL), and pit stability product, (i · x) sf , are the most influential factors. When varying chloride concentrations (RH), changes in WL, the brine specific cathodic kinetics on the external surface (captured in the equivalent current density (i eq )), and conductivity (k o ) are the most influential parameters. Finally, it was noted that dehydration reactions coupled with precipitation in the cathode will have the largest effect on predicted pit size, and cause the most significant inhibition of corrosion damage.

54 ENVIRONMENTAL SCIENCES↗

A predictive discrete-continuum multiscale model of plasticity with quantified uncertainty

Multiscale models of materials, consisting of upscaling discrete simulations to continuum models, are unique in their capability to simulate complex materials behavior. The fundamental limitation in multiscale models is the presence of uncertainty in the computational predictions delivered by them. In this work, a sequential multiscale model has been developed, incorporating discrete dislocation dynamics (DDD) simulations and a strain gradient plasticity (SGP) model to predict the size effect in plastic deformations of metallic micro-pillars. The DDD simulations include uniaxial compression of micro-pillars with different sizes and over a wide range of initial dislocation densities and spatial distributions of dislocations. An SGP model is employed at the continuum level that accounts for the size-dependency of flow stress and hardening rate. Sequences of uncertainty analyses have been performed to assess the predictive capability of the multiscale model. The variance-based global sensitivity analysis determines the effect of parameter uncertainty on the SGP model prediction. The multiscale model is then constructed by calibrating the continuum model using the data furnished by the DDD simulations. A Bayesian calibration method is implemented to quantify the uncertainty due to microstructural randomness in discrete dislocation simulations (density and spatial distribution of dislocations) on the macroscopic continuum model prediction (size effect in plastic deformation). Here, the outcomes of this study indicate that the discrete-continuum multiscale model can accurately simulate the plastic deformation of micro-pillars, despite the significant uncertainty in the DDD results. Additionally, depending on the macroscopic features represented by the DDD simulations, the SGP model can reliably predict the size effect in plasticity responses of the micropillars with below 10% of error.

36 MATERIALS SCIENCE↗

An accelerated framework for predicting creep rupture lifetimes in engineering alloys

Confidently predicting high-temperature deformation, including creep and creep rupture, is paramount for the design and commercialization of candidate materials for advanced nuclear energy systems. To accelerate creep quantification, we introduce a framework that enables rapid, cost-effective, and reliable prediction of creep rupture lifetimes, minimizing reliance on time-intensive bulk creep testing. Unlike conventional creep analysis, which requires extensive time and resources, our method leverages a maximum of four short-term bulk creep tests as training data for prediction. This framework combines high-throughput nanoindentation up to 700 °C with these targeted bulk tests to inform our creep rupture model in order to predict rupture lifetimes. The strong agreement between our predictions and conventional experimental data demonstrates the effectiveness of our approach for accelerated creep analysis and lifetime prediction of structural components in high-temperature applications. Our multi-pronged approach motivates further integration of computational tools and advanced instrumentation to establish a universal framework for understanding high-temperature material responses.

36 MATERIALS SCIENCE↗

Predicting Device Parameters for Dye-Sensitized Solar Cells from Electronic Structure Calculations to Reproduce Experiment.

Given that improvements to the power-conversion efficiency (PCE) of dye-sensitized solar cells (DSSCs) have slowed in recent years, a means to accurately predict the device parameters yielded by trial dyes in silico, without having to synthesize them, would be extremely valuable to speed up the design process. Currently, the best-performing methods of calculating device parameters rely on a set of experimentally determined kinetic coefficients. In practice, it is very difficult to measure these kinetic parameters accurately, limiting the overall accuracy of such predictive methods. This work proposes a model to obtain key parameters such as J(SC), V-OC, and PCE using only the results from density functional theory (DFT) and time-dependent DFT calculations, noting that rates of electron-transfer steps are ultimately linked to the electronic structure of the dye center dot center dot center dot TiO2 working electrode. Six organic DSSC dyes from dissimilar chemical classes (L0, L1, L2, WS-2, WS-92, and C281) were chosen to demonstrate the power of this approach. Their a priori known experimentally determined device performance metrics served to validate our predictions. The greatest absolute error in our predicted PCE values was 0.36% relative to the experiment, while the greatest fractional error was 0.042. This indicates that the proposed model offers a dramatic improvement on previous predictive methods for DSSC device parameters, both in accuracy and in consistency. Moreover, such a predictive model has great potential to be applied to other photovoltaic applications, further enabling the design of novel, highly efficient photoactive materials.

density functional theory↗

A species’ response to spatial climatic variation does not predict its response to climate change

The dominant paradigm for assessing ecological responses to climate change assumes that future states of individuals and populations can be predicted by current, species-wide performance variation across spatial climatic gradients. However, if the fates of ecological systems are better predicted by past responses to in situ climatic variation through time, this current analytical paradigm may be severely misleading. Empirically testing whether spatial or temporal climate responses better predict how species respond to climate change has been elusive, largely due to restrictive data requirements. Here, we leverage a newly collected network of ponderosa pine tree-ring time series to test whether statistically inferred responses to spatial versus temporal climatic variation better predict how trees have responded to recent climate change. When compared to observed tree growth responses to climate change since 1980, predictions derived from spatial climatic variation were wrong in both magnitude and direction. This was not the case for predictions derived from climatic variation through time, which were able to replicate observed responses well. Future climate scenarios through the end of the 21st century exacerbated these disparities. These results suggest that the currently dominant paradigm of forecasting the ecological impacts of climate change based on spatial climatic variation may be severely misleading over decadal to centennial timescales.

54 ENVIRONMENTAL SCIENCES↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic prediction of switchgrass winter survivorship across diverse lowland populations

Abstract In the North-Central United States, lowland ecotype switchgrass can increase yield by up to 50% compared with locally adapted but early flowering cultivars. However, lowland ecotypes are not winter tolerant. The mechanism for winter damage is unknown but previously has been associated with late flowering time. This study investigated heading date (measured for two years) and winter survivorship (measured for three years) in a multi-generation population generated from two winter-hardy lowland individuals and diverse southern lowland populations. Sequencing data (311,776 markers) from 1,306 individuals were used to evaluate genome-wide trait prediction through cross-validation and progeny prediction (n = 52). Genetic variance for heading date and winter survivorship was additive with high narrow-sense heritability (0.64 and 0.71, respectively) and reliability (0.68 and 0.76, respectively). The initial negative correlation between winter survivorship and heading date degraded across generations (F1 r = −0.43, pseudo-F2 r = −0.28, pseudo-F2 progeny r = −0.15). Within-family predictive ability was moderately high for heading date and winter survivorship (0.53 and 0.52, respectively). A multi-trait model did not improve predictive ability for either trait. Progeny predictive ability was 0.71 for winter survivorship and 0.53 for heading date. These results suggest that lowland ecotype populations can obtain sufficient survival rates in the northern United States with two or three cycles of effective selection. Despite accurate genomic prediction, naturally occurring winter mortality successfully isolated winter tolerant genotypes and appears to be an efficient method to develop high-yielding, cold-tolerant switchgrass cultivars.

60 APPLIED LIFE SCIENCES↗

Genomic prediction of regional-scale performance in switchgrass ( Panicum virgatum ) by accounting for genotype-by-environment variation and yield surrogate traits

Switchgrass is a potential crop for bioenergy or carbon capture schemes, but further yield improvements through selective breeding are needed to encourage commercialization. To identify promising switchgrass germplasm for future breeding efforts, we conducted multisite and multitrait genomic prediction with a diversity panel of 630 genotypes from 4 switchgrass subpopulations (Gulf, Midwest, Coastal, and Texas), which were measured for spaced plant biomass yield across 10 sites. Our study focused on the use of genomic prediction to share information among traits and environments. Specifically, we evaluated the predictive ability of cross-validation (CV) schemes using only genetic data and the training set (cross-validation 1: CV1), a subset of the sites (cross-validation 2: CV2), and/or with 2 yield surrogates (flowering time and fall plant height). We found that genotype-by-environment interactions were largely due to the north–south distribution of sites. The genetic correlations between the yield surrogates and the biomass yield were generally positive (mean height r = 0.85; mean flowering time r = 0.45) and did not vary due to subpopulation or growing region (North, Middle, or South). Genomic prediction models had CV predictive abilities of –0.02 for individuals using only genetic data (CV1), but 0.55, 0.69, 0.76, 0.81, and 0.84 for individuals with biomass performance data from 1, 2, 3, 4, and 5 sites included in the training data (CV2), respectively. To simulate a resource-limited breeding program, we determined the predictive ability of models provided with the following: 1 site observation of flowering time (0.39); 1 site observation of flowering time and fall height (0.51); 1 site observation of fall height (0.52); 1 site observation of biomass (0.55); and 5 site observations of biomass yield (0.84). The ability to share information at a regional scale is very encouraging, but further research is required to accurately translate spaced plant biomass to commercial-scale sward biomass performance.

09 BIOMASS FUELS↗

Optimizing genomic prediction for complex traits via investigating multiple factors in switchgrass

Genomic prediction has accelerated breeding processes and provided mechanistic insights into the genetic bases of complex traits. To further optimize genomic prediction, we assess the impact of genome assemblies, genotyping approaches, variant types, allelic complexities, polyploidy levels, and population structures on the prediction of 20 complex traits in switchgrass (Panicum virgatum L.), a perennial biofuel feedstock. Surprisingly, short read-based genome assembly performs comparably to or even better than long read-based assembly. Due to higher gene coverage, exome capture and multi-allelic variants outperform genotyping-by-sequencing and bi-allelic variants, respectively. Tetraploid models show higher prediction accuracy than octoploid models for most traits, likely due to the greater genetic distances among tetraploids. Depending on the trait in question, different types of variants need to be integrated for optimal predictions. Furthermore, our study provides insights into the factors influencing genomic prediction outcomes, guiding best practices for future studies and for improving agronomic traits in switchgrass and other species through selective breeding.

60 APPLIED LIFE SCIENCES↗

Machine Learning-Based Prediction of Distribution Network Voltage and Sensor Allocation

Increasing penetration levels of fast-varying energy resources might negatively affect power system operation. At the same time, sensor deployment throughout distribution networks improves system awareness and enables the development of new and advanced voltage control solutions. Such control techniques rely on accurate prediction in anticipation of voltage violation scenarios. This paper analyzes various approaches to voltage prediction in a distribution system, and it is shown that combining multiple techniques into a single regressor improves its predictive power. Moreover, a two-step regressor is proposed in which initial predictions based on a global regressor are refined by local regressors; in this case, prediction errors decrease significantly. Additionally, a clustering approach is employed to perform sensor allocation so that only the most influential buses are selected for monitoring without diminishing prediction accuracy.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neural-Network-Enhanced COTSIM: Advancing Predictive Capabilities for Fast DIII-D Simulations

Sustaining fusion reactions in tokamaks requires heating plasma to thermonuclear temperatures while maintaining confinement and stability. Neutral beam injection (NBI) provides heating, current drive, torque, and fueling, while electron cyclotron (EC) waves are widely used for heating and current drive; together, these actuators shape the plasma current, temperature, and density profiles. The control-oriented tokamak simulator (COTSIM), a predictive, control-oriented code, has been enhanced with neural-network surrogates for transport and sources. Turbulent transport is predicted by MMMnet—a neural-network version of the updated multimode model (MMM 9.0.10)—with significantly reduced computation time relative to MMM; neoclassical transport follows the Chang–Hinton model. NUBEAMnet, a surrogate of the Monte Carlo NUBEAM module, predicts beam-driven heating, current, and torque. EC heating and current drive use a control-oriented, empirically scaled source model; plasma resistivity follows the Spitzer formulation; bootstrap current uses the Sauter model. Equilibrium is computed using both prescribed and fixed-boundary solvers (FBSs), and the pedestal structure is modeled with an empirical pedestal model. For a representative DIII-D discharge, COTSIM predicts electron and ion temperature and safety-factor profiles in close agreement with TRANSP predictive and interpretive simulations while extending predictions through the pedestal region to the plasma edge (versus 80% of the minor radius in TRANSP). Furthermore, the equivalent COTSIM simulation runs in under 3 min compared to about 2 h for TRANSP, enabling rapid scenario planning, optimization of tokamak operation, and between-pulse control design.

Control-oriented tokamak simulator (COTSIM)↗

Machine Learning-enabled Scalable Performance Prediction of Scientific Codes

Hardware architectures become increasingly complex as the compute capabilities grow to exascale. Here, we present the Analytical Memory Model with Pipelines (AMMP) of the Performance Prediction Toolkit (PPT). PPT-AMMP takes high-level source code and hardware architecture parameters as input and predicts runtime of that code on the target hardware platform, which is defined in the input parameters. PPT-AMMP transforms the code to an (architecture-independent) intermediate representation, then (i) analyzes the basic block structure of the code, (ii) processes architecture-independent virtual memory access patterns that it uses to build memory reuse distance distribution models for each basic block, and (iii) runs detailed basic-block level simulations to determine hardware pipeline usage. PPT-AMMP uses machine learning and regression techniques to build the prediction models based on small instances of the input code, then integrates into a higher-order discrete-event simulation model of PPT running on Simian PDES engine. We validate PPT-AMMP on four standard computational physics benchmarks and present a use case of hardware parameter sensitivity analysis to identify bottleneck hardware resources on different code inputs. We further extend PPT-AMMP to predict the performance of a scientific application code, namely, the radiation transport mini-app SNAP. To this end, we analyze multi-variate regression models that accurately predict the reuse profiles and the basic block counts. We validate predicted SNAP runtimes against actual measured times.

97 MATHEMATICS AND COMPUTING↗

Quantifying Uncertainty in HPC Job Queue Time Predictions

High Performance Computing (HPC) has developed at an unprecedented pace in recent decades. This growth has demanded corresponding development in the area of HPC Operational Data Analytics (ODA), which encompasses a wide range of data analysis techniques, ML/AI efforts, tools, and visualizations. Published studies in ODA offer a variety of practical ways to inform HPC users, administrators, procurement managers, and other stakeholders. Uncertainty analysis, however, is rare in the related published literature. For instance, we identify only 1 out of 14 existing studies focused on job queue time prediction that investigates the uncertainty aspect of their proposed predictions. We recognize the utmost importance uncertainty quantification can have in such predictive analytics solutions, with consequences in how users interpret information they receive, and attempt to bridge this gap. With the goal of improving access to such insights, we develop a process for determining upper and lower bounds of the predicted queue times of a regression model at a specified confidence level. Our current research is focused on the uncertainty in predicting job queue times, yet our approach may be employed in predicting other metrics.

HPC↗

Comparing Individualized Survival Predictions From Random Survival Forests and Multistate Models in the Presence of Missing Data: A Case Study of Patients With Oropharyngeal Cancer

Background: In recent years, interest in prognostic calculators for predicting patient health outcomes has grown with the popularity of personalized medicine. These calculators, which can inform treatment decisions, employ many different methods, each of which has advantages and disadvantages. Methods: We present a comparison of a multistate model (MSM) and a random survival forest (RSF) through a case study of prognostic predictions for patients with oropharyngeal squamous cell carcinoma. The MSM is highly structured and takes into account some aspects of the clinical context and knowledge about oropharyngeal cancer, while the RSF can be thought of as a black-box non-parametric approach. Key in this comparison are the high rate of missing values within these data and the different approaches used by the MSM and RSF to handle missingness. Results: We compare the accuracy (discrimination and calibration) of survival probabilities predicted by both approaches and use simulation studies to better understand how predictive accuracy is influenced by the approach to (1) handling missing data and (2) modeling structural/disease progression information present in the data. We conclude that both approaches have similar predictive accuracy, with a slight advantage going to the MSM. Conclusions: Although the MSM shows slightly better predictive ability than the RSF, consideration of other differences are key when selecting the best approach for addressing a specific research question. These key differences include the methods’ ability to incorporate domain knowledge, and their ability to handle missing data as well as their interpretability, and ease of implementation. Ultimately, selecting the statistical method that has the most potential to aid in clinical decisions requires thoughtful consideration of the specific goals.

60 APPLIED LIFE SCIENCES↗

DeepGRN: prediction of transcription factor binding site across cell-types using attention-based deep neural networks

Abstract Background Due to the complexity of the biological systems, the prediction of the potential DNA binding sites for transcription factors remains a difficult problem in computational biology. Genomic DNA sequences and experimental results from parallel sequencing provide available information about the affinity and accessibility of genome and are commonly used features in binding sites prediction. The attention mechanism in deep learning has shown its capability to learn long-range dependencies from sequential data, such as sentences and voices. Until now, no study has applied this approach in binding site inference from massively parallel sequencing data. The successful applications of attention mechanism in similar input contexts motivate us to build and test new methods that can accurately determine the binding sites of transcription factors. Results In this study, we propose a novel tool (named DeepGRN) for transcription factors binding site prediction based on the combination of two components: single attention module and pairwise attention module. The performance of our methods is evaluated on the ENCODE-DREAM in vivo Transcription Factor Binding Site Prediction Challenge datasets. The results show that DeepGRN achieves higher unified scores in 6 of 13 targets than any of the top four methods in the DREAM challenge. We also demonstrate that the attention weights learned by the model are correlated with potential informative inputs, such as DNase-Seq coverage and motifs, which provide possible explanations for the predictive improvements in DeepGRN. Conclusions DeepGRN can automatically and effectively predict transcription factor binding sites from DNA sequences and DNase-Seq coverage. Furthermore, the visualization techniques we developed for the attention modules help to interpret how critical patterns from different types of input features are recognized by our model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗

Subsurface Characterization and Machine Learning Predictions at Brady Hot Springs Results

Geothermal power plants typically show decreasing heat and power production rates over time. Mitigation strategies include optimizing the management of existing wells - increasing or decreasing the fluid flow rates across the wells - and drilling new wells at appropriate locations. The latter is expensive, time-consuming, and subject to many engineering constraints, but the former is a viable mechanism for periodic adjustment of the available fluid allocations. Data and supporting literature from a study describing a new approach combining reservoir modeling and machine learning to produce models that enable strategies for the mitigation of decreased heat and power production rates over time for geothermal power plants. The computational approach used enables translation of sets of potential flow rates for the active wells into reservoir-wide estimates of produced energy and discovery of optimal flow allocations among the studied sets. In our computational experiments, we utilize collections of simulations for a specific reservoir (which capture subsurface characterization and realize history matching) along with machine learning models that predict temperature and pressure timeseries for production wells. We evaluate this approach using an "open-source" reservoir we have constructed that captures many of the characteristics of Brady Hot Springs, a commercially operational geothermal field in Nevada, USA. Selected results from a reservoir model of Brady Hot Springs itself are presented to show successful application to an existing system. In both cases, energy predictions prove to be highly accurate: all observed prediction errors do not exceed 3.68% for temperatures and 4.75% for pressures. In a cumulative energy estimation, we observe prediction errors that are less than 4.04%. A typical reservoir simulation for Brady Hot Springs completes in approximately 4 hours, whereas our machine learning models yield accurate 20-year predictions for temperatures, pressures, and produced energy in 0.9 seconds. This paper aims to demonstrate how the models and techniques from our study can be applied to achieve rapid exploration of controlled parameters and optimization of other geothermal reservoirs. Includes a synthetic, yet realistic, model of a geothermal reservoir, referred to as open-source reservoir (OSR). OSR is a 10-well (4 injection wells and 6 production wells) system that resembles Brady Hot Springs (a commercially operational geothermal field in Nevada, USA) at a high level but has a number of sufficiently modified characteristics (which renders any possible similarity between specific characteristics like temperatures and pressures as purely random). We study OSR through CMG simulations with a wide range of flow allocation scenarios. Includes a dataset with 101 simulated scenarios that cover the period of time between 2020 and 2040 and a link to the published paper about this project, where we focus on the Machine Learning work for predicting OSR's energy production based on the simulation data, as well as a link to the GitHub repository where we have published the code we have developed (please refer to the repository's readme file to see instructions on how to run the code). Additional links are included to associated work led by the USGS to identify geologic factors associated with well productivity in geothermal fields. Below are the high-level steps for applying the same modeling + ML process to other geothermal reservoirs: 1. Develop a geologic model of the geothermal field. The location of faults, upflow zones, aquifers, etc. need to be accounted for as accurately as possible 2. The geologic model needs to be converted to a reservoir model that can be used in a reservoir simulator, such as, for instance, CMG STARS, TETRAD, or FALCON 3. Using native state modeling, the initial temperature and pressure distributions are evaluated, and they become the initial conditions for dynamic reservoir simulations 4....

15 GEOTHERMAL ENERGY↗

Predicting Execution Times for Disk-based and In-Situ Parallel Data Analytics (Final Technical Report)

In recent years, there has been a significant amount of interests in in-situ analytics on simulation programs. For a variety of reasons, it is desirable to be able to predict the execution time of an analytics program. At the same time, frameworks such as MapReduce have become popular for scientific data analytics. This paper focuses on developing performance models for predicting execution time of parallel data analytics, with a special emphasis on in-situ analytics. We take two distinct approach towards performance prediction. We first expand SKOPE (a SKeleton framewOrk for Performance Exploration) with performance models for disk data read, cache performance, and page fault penalty. Second, an analytical performance model is also developed. We have evaluated our performance prediction framework as well as the analytical model on three hardware setups with well-known data mining algorithms implemented in three programming paradigms, MapReduce, MATE (a MapReduce-like parallel system with an alternate API for multi-core environments) and Smart (a MapReduce-like framework for in-situ analytics). Results show that our performance prediction framework along with the incorporated performance models are capable of accurately predicting execution times for parallel scientific analytics on different hardware setups.

97 MATHEMATICS AND COMPUTING↗