Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data exploration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

DELVE-ing into the Milky Way’s Globular Clusters: Assessing Extratidal Features in NGC 5897, NGC 7492, and Testing Detectability with Deeper Photometry

Extratidal features around globular clusters (GCs) are tracers of their disruption, stellar stream formation, and their host’s gravitational potential. However, these features remain challenging to detect due to their low surface brightness. We conduct a systematic search for such features around 19 GCs in the DECam Local Volume Exploration (DELVE) survey Data Release 2, discovering a new extra-tidal envelope around NGC 5897 and find tentative evidence for an extended envelope surrounding NGC 7492. Through a combination of dynamical modeling and analyzing synthetic stellar populations, we demonstrate these envelopes may have formed through tidal disruption. We use these models to explore the detectability of these features in the upcoming Legacy Survey of Space and Time (LSST), finding that while LSST’s deeper photometry will enhance detection significance, additional methods for foreground removal like proper motions or metallicities may be important for robust stream detection. Our results both add to the sample of globular clusters with extratidal features and provide insights on interpreting similar features in current and upcoming data.

Chiti, A. [Univ. of Chicago, IL (United States); S↗

Ducted Fuel Injection And Cooled Spray Technologies For Particulate Control In Heavy-duty Diesel Engines (Final Report)

Cooled Spray (CS) and Ducted Fuel Injection (DFI) are in-cylinder technologies for diesel engines that can reduce particulate matter and soot emissions and data has been published showing that these technologies can reduce soot emissions by 75-100% for some engines at some operating conditions. However, little is known about scaling the devices for engine size. Additionally, the performance of either technology over the engine duty cycle has not been explored. This project addresses both of these points through single-cylinder engine investigations. The objectives of this project are to provide details about dimensional scaling of these devices and to demonstrate 75% PM reduction over a range of operating conditions on a single-cylinder engine. Two engines were used for this project: a 125mm bore optically accessible engine at Sandia National Laboratories and a 168mm bore metal engine at Southwest Research Institute. The optical engine was used to study the performance of DFI and CS inserts for a large injector orifice diameter injector that is characteristic of a locomotive engine and to perform scaling studies for DFI. The metal engine was used to perform scaling and alignment studies for CS and to evaluate the technology for both EGR and non-EGR engines over the engine operating map. Modifications were required for both engines to accept the prototype inserts being tested. The optical engine required a new fuel injector, cylinder head and piston so that tests could be run at the pressures and engine speeds required. Additionally, a novel rotating stage was designed for the optical engine to simplify alignment of the modules. The metal engine required a modified cylinder head to accept CS inserts and a modified piston to provide additional space around the fuel injector for the CS inserts. Tests on the optical engine showed that DFI reduces PM emissions for both small injector orifices (0.170mm diameter) and large injector orifices (0.290mm). For high load testing, the DFI modules were not as effective as at low load testing, but it was acknowledged that minimal geometric optimization was performed and more improvements may be possible. Comparing DFI to CS and conventional diesel combustion (CDC), DFI performed better than CS or CDC. The CS geometries used in these studies may not be ideal for that engine and additional modifications likely would improve performance. Tests on the metal engine showed PM reductions as high has 80% at some operating conditions with duty-cycle PM reductions of ~50% for EGR and non-EGR configurations. The CS testing on the metal engine showed that chamfering of the fuel passage inlet either through hydro-erosion or mechanical grinding provided significant improvements in the PM reduction capabilities of the insert. Additionally, alignment sensitivities were explored and the data show that the tolerance to misalignment is approximately 0.05 to 0.1mm for the inserts that were studied here. Air-fuel ratio was shown to be important in the effectiveness of the CS inserts. In several tests, it was shown that the CS inserts are more effective at reducing the PM for high-AFR operating conditions compared to low AFR conditions. In summary, multiple designs were evaluated on both engines. It was found that for the conditions and configurations studied here, a fuel passage diameter of ~2.5mm performed best overall. Significant duty-cycle PM reductions are possible using these technologies and sensitivities to AFR, alignment fuel passage diameter and inlet fuel passage shaping were explored and are reported here. More PM reduction may be possible with improved geometric design and attention to alignment practices.

02 PETROLEUM↗

An idea to explore: How an interdisciplinary undergraduate course exploring a global health challenge in molecular detail enabled science communication and collaboration in diverse audiences

Abstract Communication and collaboration are key science competencies that support sharing of scientific knowledge with experts and non‐experts alike. On the one hand, they facilitate interdisciplinary conversations between students, educators, and researchers, while on the other they improve public awareness, enable informed choices, and impact policy decisions. Herein, we describe an interdisciplinary undergraduate course focused on using data from various bioinformatics data resources to explore the molecular underpinnings of diabetes mellitus (Types 1 and 2) and introducing students to science communication. Building on course materials and original student‐generated artifacts, a series of collaborative activities engaged students, educators, researchers, healthcare professionals and community members in exploring, learning about, and discussing the molecular bases of diabetes. These collaborations generated novel educational materials and approaches to learning and presenting complex ideas about major global health challenges in formats accessible to diverse audiences.

59 BASIC BIOLOGICAL SCIENCES↗

Cell‐type‐specific transcriptomics uncovers spatial regulatory networks in bioenergy sorghum stems

SUMMARY Bioenergy sorghum is a low‐input, drought‐resilient, deep‐rooting annual crop that has high biomass yield potential enabling the sustainable production of biofuels, biopower, and bioproducts. Bioenergy sorghum's 4–5 m stems account for ~80% of the harvested biomass. Stems accumulate high levels of sucrose that could be used to synthesize bioethanol and useful biopolymers if information about cell‐type gene expression and regulation in stems was available to enable engineering. To obtain this information, laser capture microdissection was used to isolate and collect transcriptome profiles from five major cell types that are present in stems of the sweet sorghum Wray. Transcriptome analysis identified genes with cell‐type‐specific and cell‐preferred expression patterns that reflect the distinct metabolic, transport, and regulatory functions of each cell type. Analysis of cell‐type‐specific gene regulatory networks (GRNs) revealed that unique transcription factor families contribute to distinct regulatory landscapes, where regulation is organized through various modes and identifiable network motifs. Cell‐specific transcriptome data was combined with known secondary cell wall (SCW) networks to identify the GRNs that differentially activate SCW formation in vascular sclerenchyma and epidermal cells. The spatial transcriptomic dataset provides a valuable source of information about the function of different sorghum cell types and GRNs that will enable the engineering of bioenergy sorghum stems, and an interactive web application developed during this project will allow easy access and exploration of the data ( https://mc‐lab.shinyapps.io/lcm‐dataset/ ).

09 BIOMASS FUELS↗

Scenario Discovery Analysis of Drivers of Solar and Wind Energy Transitions Through 2050

Deep human-Earth system uncertainties and strong multi-sector dynamics make it difficult to anticipate which conditions are most likely to lead to higher or lower adoption of renewable energy, and models project a broad range of future solar and wind energy shares across future scenarios. To elucidate these dynamics, we explore a large data set of scenarios simulated from the Global Change Analysis Model (GCAM) and use scenario discovery to identify the most significant factors affecting solar and wind adoption by mid-century. We generated a data set of over 4,000 scenarios from GCAM by varying 12 different socioeconomic factors at high and low levels, including assumptions about future energy demand, resource costs, and fossil fuel emissions paths, as well as specific technology assumptions including wind and solar backup requirements and storage costs. Using scenario discovery, we assess the most important factors globally and regionally in creating high fractions of solar and wind energy and explore interconnected effects on other systems including water and non-CO 2 emissions. Globally and regionally, we found that solar and wind-related technology costs were the primary drivers of high wind and solar energy adoption, though a few regions depend heavily on other parameters like carbon capture and storage costs, population and gross domestic product trajectories, and fossil fuel costs. We also identify four key paths to high solar and wind energy by mid-century and discuss their tradeoffs in terms of other outcomes.

14 SOLAR ENERGY↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Methodology for physics-informed generation of synthetic neutron time-of-flight measurement data

Accurate neutron cross section data are a vital input to the simulation of nuclear systems for a wide range of applications from energy production to national security. The evaluation of experimental data is a key step in producing accurate cross sections. There is a widely recognized lack of reproducibility in the evaluation process due to its artisanal nature and therefore there is a call for improvement within the nuclear data community. This can be realized by automating/standardizing viable parts of the process, namely, parameter estimation by fitting theoretical models to experimental data. This automation effort could greatly benefit from a synthetic data resource. This work leverages problem-specific physics, Monte Carlo sampling, and a general methodology for data synthesis to generate unlimited, labelled experimental cross-section data that is statistically indistinguishable to the observed data. Heuristic and, where applicable, rigorous statistical comparisons to observed data support this claim. The demonstration is based on/limited to transmission measurements at Rensselaer Polytechnic Institute (RPI) and energy-differential cross sections in the resolved resonance region (RRR). An open-source software is published alongside this article that executes the complete methodology to produce high-utility synthetic datasets. The goal of this work is to provide an approach and corresponding tool that will allow the evaluation community to begin exploring more data-driven, ML-based solutions to long-standing challenges in the field.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Developing Novel Performance Measures for Traffic Congestion Management and Operational Planning Based on Connected Vehicle Data

In this study, the authors present their efforts in exploring a new type of traffic data, referred to as internet-connected vehicle (ICV) data, for traffic congestion management and operational planning. Most currently manufactured vehicles contain onboard GPS and cellular modules, and they constantly connect to automobile manufacturers' clouds via cellular networks and upload their status. Some automobile manufacturers have recently redistributed the nonpersonal part of such data, such as geolocation, to third-party organizations for innovative applications. Compared with the traditional vehicle GPS data, the ICV data contain high-resolution GPS waypoints accompanied with the vehicles' abnormal moving events (e.g., hard braking). The ICV data also have huge potential in congestion management and operational planning. They explore to identify and analyze traffic congestion on both freeways and arterials using the ICV data. The ICV data adopted for this research are redistributed by Wejo Data Service, representing 10%-15% of all moving vehicles in the Dallas-Fort Worth (DFW) area in Texas. Through one case study for a freeway segment and one for an arterial segment, new traffic performance metrics based on the characteristics of ICV data have been presented. The highlights of these efforts are as follows: (I) queue length and propagation at freeway bottlenecks can be directly measured based on where and when most internet-connected vehicles slow down and join the queue; (II) an internet-connected vehicle's actual delay time on arterials can be directly measured according to its slow movement percentage, without assuming the nondelay travel speed; and (III) the ICV data set are also combined with the high-resolution traffic signal events to generate a ground-truth time-space diagram (TSD) on arterials - a common visualization of arterial signal performance for transportation planning and operations.

33 ADVANCED PROPULSION SYSTEMS↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

First-Principles Study on the Electronic Properties of PDPP-Based Conjugated Polymer via Density Functional Theory

In this study, we focus on computational predictions of the electronic and optical properties of a one-dimensional periodic model of a single chain of a diketopyrrolopyrrole (DPP)-based conjugated polymer (PDPP3T) as a function of electronic configuration changes due to charge injection. Here, we employ density functional theory (DFT) to explore the ground-state and excited-state electronic properties as well as optical properties influenced by charge injection. We utilize both the Heyd–Scuseria–Ernzerhof (HSE06) and Perdew–Burke–Ernzerhof (PBE) functionals to predict the band gap and compute the absorption spectrum. Our DFT results point out that utilizing the HSE06 functional in conjunction with momentum sampling over the Brillouin zone can appropriately predict the band gap and absorption spectrum in good agreement with experimental data. Moreover, we explore the influence of charge-carrier injection on the electronic configuration of the PDPP3T polymer. Our results indicate that the injection of charge carriers into the PDPP3T semiconducting polymer model greatly affects the electrical properties and ends in a low band gap and high mobility of charge carriers in PDPP3T polymers, offering the potential to tailor the material electronic performance for organic photovoltaic and optoelectronic device applications.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Novel machine-learning method for spin classification of neutron resonances

The performance of nuclear reactors and other nuclear systems depends on a precise understanding of the neutron interaction cross sections for materials used in these systems. These cross sections exhibit resonant structure whose shape is determined in part by the angular-momentum quantum numbers of the resonances. The correct assignment of the quantum numbers of neutron resonances is, therefore, paramount. In this project, we apply machine learning to automate the quantum number assignments using only the resonances' energies and widths and not relying on detailed transmission or capture measurements. The classifier used for quantum number assignment is trained using stochastically generated resonance sequences whose distributions mimic those of real data. Here we explore the use of several physics-motivated features for training our classifier. These features amount to out-of-distribution tests of a given resonance's widths and resonance-pair spacings. We pay special attention to situations where either capture widths cannot be trusted for classification purposes or where there is insufficient information to classify resonances by the total spin J. We demonstrate the efficacy of our classification approach using simulated and actual 52 Cr resonance data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

Utility-Scale Solar, 2024 Edition: Empirical Trends in Deployment, Technology, Cost, Performance, PPA Pricing, and Value in the United States [Slides]

Berkeley Lab’s “Utility-Scale Solar, 2024 Edition” presents analysis of empirical plant-level data from the U.S. fleet of ground-mounted photovoltaic (PV), PV+battery, and concentrating solar-thermal power (CSP) plants with capacities exceeding 5 MWAC (PV plants of 5 MWAC or less, including residential rooftop systems, are covered separately in Berkeley Lab’s companion annual report, Tracking the Sun). Key findings from this year’s report include: -18.5 GWAC of new utility-scale PV capacity came online in 2023, bringing cumulative installed capacity to more than 80.2 GWAC across 47 states. Installed costs continued to fall in 2023. Relative to 2022, capacity-weighted averages decreased by 8% to -$\$1.43$/WAC (or $\$1.08$/WDC). Costs, based on a 7.1 GWAC sample of 76 plants completed in 2023, have fallen by 75% (averaging 10% annually) since 2010. Plant-level capacity factors vary widely, from 6% to 36% (on an AC basis), with a sample median of 24%. -Levelized cost of energy (LCOE) of new 2023 projects increased slightly to $\$46$/MWh prior to the application of tax credits but continued to fall to $\$31$/MWh when accounting for federal incentives. PPA prices have largely followed the decline in solar’s LCOE over time, but newly signed longer-term PPA prices have increased since 2021, to an average of $\$35$/MWh (levelized, in 2023 dollars). -Solar’s average energy and capacity value (i.e., ability to offset costs of other power generation sources) across the U.S. was $\$45$/MWh in 2023. Solar’s average market value was lowest in CAISO ($\$27$/MWh), the market with the greatest solar generation share, and highest in ERCOT ($\$67$/MWh). -Newer solar projects had greater market value in 2023 than their generation costs, yielding $\$1.1$ billion in benefits. Projects built in 2022 delivered on average $\$15$/MWh more market value than their costs in 2023. -Solar’s combined value from wholesale electricity markets, public health and climate damage reduction were greater than generation costs and incentives, yielding $\$13.7$ billion in net benefits in 2023. We estimate U.S. health benefits of $\$24$/MWh and reduced global climate damages of $\$101$/MWh. -Adding battery storage is one way to increase the value of solar. Deployment of 52 new PV+battery hybrid plants set a record with 5.3 GW installed in 2023. Our public data file tracks metadata and PPA prices from more than 100 PV+battery hybrid projects that are already online or that have secured offtake arrangements. -Looking ahead, a massive pipeline of at least 1,085 GW of solar capacity dominates the nation’s interconnection queues at the end of 2023. Nearly 571 GW, or 53%, of that total was paired with a battery – in CAISO it was a staggering 98%. Historically only 10% of the requested solar capacity is built. -For more information, and to explore related interactive data visualizations, go to utilityscalesolar.lbl.gov.

14 SOLAR ENERGY↗

Distribution Development for the RDX Regional Model at Los Alamos National Laboratory - 20385

Representing uncertainty in model inputs often means finding a balance between uncertainty and physical reality. Developing wide distributions may seem conservative in principle, but this approach may lead to unrealistic model results. Characterizing the current state of knowledge of stochastic inputs presents many challenges, especially if the data available are limited or have limited relevance to the site. Relationships among these inputs may also be important to represent but are typically complex or difficult to define. Often special adjustments must be made to account for reduced credibility in particular data. If parameters are strongly related to one another, a correlation structure may be developed for input to the model. Other techniques such as regression models may be used to incorporate relationships between the information available and the desired parameters. The process of developing distributions must consider details of the model in terms of what the distribution is meant to represent. This paper uses the example of a probabilistic fate and transport model for hexahydro-1,3,5-trinitro-1,3,5-triazine (RDX) in the regional aquifer at Los Alamos National Laboratory (LANL). For many parameters, a single draw is applied to all space and time over which the model is run, for a single iteration. This simplification is often made for many reasons, and can often be beneficial, but also adds additional complexity in the distribution development process. Defining the distributional goals as they relate to the modeling process is an important step, which should take place prior to evaluation of the data. Usually, the distributions developed are meant to characterize the average value of the parameter over the spatial and temporal domain of the model. Distribution development requires consideration of many sources of information on the parameter where available, ideally from multiple references. Examples of different sources include data from different references but also from different conditions, measurement methods, or experimental types. Depending on these conditions and the reliability or relevance of particular references, different sources of data may each contribute valuable information but have varying relevance to the site. In these cases, weighting data unequally is a useful way to incorporate this information. As an example, aqueous dispersivity data are available for a variety of rock types. Only a few values are available for the desired rock type, and this is not enough to develop a distribution. Therefore, dispersivity values from other rock types are included in distribution development but are down-weighted such that the best data have the most influence on the distribution developed. In another case, K{sub d} distributions in the model are meant to represent a known composition of multiple soil types. Data from these materials are weighted accordingly to develop a distribution for the weighted average K{sub d} across soil types. Other cases include varying reliability of different sources, and weighting data according to the confidence in these sources. In some cases, input parameters are correlated with one another. An example is advective porosity, which is positively correlated with total porosity and must be less than total porosity. Paired data with both parameters must exist to discern the relationship between the two parameters and if it is necessary to build a correlation structure into the model. In general, correlated parameters may be represented in the model by a multivariate distribution, or perhaps more desirably, capturing the correlation within the developed distributions which may be treated as independent from one another. In the example of porosity, this can be done by transforming advective porosity into a proportion of total porosity which may be drawn independently from the distribution of total porosity. This paper explores how complex data and correlations can be incorporated to meet the distributional goals of the model, using the RDX regional model as a detailed example. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Molecular Dynamical and Quantum Mechanical Exploration of the Site-Specific Dynamics of Cy3 Dimers Internally Linked to dsDNA

Performing spectroscopic measurements on biomolecules labeled with fluorescent probes is a powerful approach to locating the molecular behavior and dynamics of large systems at specific sites within their local environments. The indocarbocyanine dye Cy3 has emerged as one of the most commonly used chromophores. The incorporation of Cy3 dimers into DNA enhances experimental resolution owing to the spectral characteristics influenced by the geometric orientation of excitonically coupled monomeric units. Various theoretical models and simulations have been utilized to aid in the interpretation of the experimental spectra. In this study, we employ all-atom molecular dynamics simulations to study the structural dynamics of Cy3 dimers internally linked to the dsDNA backbone. We used quantum mechanical calculations to derive insights from both the linear absorption spectra and the circular dichroism data. Furthermore, we explore potential limitations within a commonly used force field for cyanine dyes. The molecular dynamics simulations suggest the presence of four possible Cy3 dimeric populations. The spectral simulations on the four populations show one of them to agree better with the experimental signatures, suggesting it to be the dominant population. Furthermore, the relative orientation of Cy3 in this population compares very well with previous predictions from the Holstein–Frenkel Hamiltonian model.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗