Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Molecular Modeling and Molecular Dynamics Simulation of a Packed and Intact Bacterial Microcompartment

Bacterial microcompartments (BMCs) are protein-bound organelles found in some bacteria which encapsulate enzymes for enhanced catalytic activity. These compartments spatially sequester enzymes within semipermeable shell proteins and are packed full of enzyme cargoes and metabolites as they fulfill their function. Coupling together recent SAXS and proteomics work, it is possible to develop molecular models for these microcompartments and interrogate enzyme and metabolite dynamics within. Our primary goal of this study is to quantify the permeability of metabolite glyceraldehyde-3-phosphate (G3P) and dihydroxyacetone phosphate (DHAP) across the BMC shell through classical molecular dynamics simulation. The Haliangium ochraceum model of BMC shell (PDB: 6MZX) was used to model an intact BMC of approximately 10 million atoms. Working at this scale presented its own challenges in managing large data sets, with multiple challenges and hardware advances discussed that facilitated this work. Over approximately 750 ns of aggregate simulation, we see multiple permeation events for these metabolites that were added at high concentration through the pores present within BMC shell tiles. When compared to independent permeability estimates for the same metabolites determined through replica exchange umbrella sampling simulations, the permeabilities varied by approximately 3 orders of magnitude. Regardless, the permeability coefficients for both G3P and DHAP are highly similar and very high, such that only very small concentration gradients can be maintained across the BMC shell between the cytosol and BMC interior. The large simulation systems also facilitated comparisons for molecular diffusivity in the crowded environment within the BMC shell. By our estimates, the viscosity within a packed BMC shell is at least 10-fold higher than it would be in neat solution and is the real driver for varying permeability estimates we obtained through simulation. These findings will be used as design inputs for future bioengineering efforts to make products from BMCs, highlighting how permeable BMC shells can be.

Diffusion↗

Generative AI for Grid Operations [Slides]

In the last few years, the development and use of generative artificial intelligence (AI) and large-language models (LLMs) have changed the landscape of how AI and machine learning (ML) are being used in power systems. LLMs are built on foundational models based on large data sets that can be trained to provide information rapidly and through simple natural language prompts. Generative AI can then perform human-like tasks using ML models to identify and mimic pattens in the data sets. This presentation explores how generative AI can enhance grid operations by improving forecasts, enabling rapid contingency analyses, and offering real-time operational suggestions. By providing grid operators with valuable insights, generative AI will empower them to manage power systems more effectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scalable probabilistic estimates of electric vehicle charging given observed driver behavior

To prepare for rapid growth in global electric vehicle adoption, grid and policy planners depend on detailed forecasts of future charging demand. In this paper we propose a novel holistic, scalable, probabilistic framework to produce large-scale estimates of electric vehicle charging load for long-term planning that capture real drivers’ charging patterns. Our framework captures the uncertainty and stochasticity in charging demand by taking a graphical modeling approach. It has three core elements: driver groups, charging segment choices, and charging session time and energy requirements. The framework uses hierarchical clustering to group drivers by their charging histories, capturing their heterogeneous behaviors and preferences across different segments or types of charging. The framework uses probabilistic mixture models for each driver group’s sessions to identify the unique charging behaviors observed within each segment. We illustrate its application with a large data set from California, profiling the charging patterns and unique driver clusters it identifies. Using the model knobs representing drivers’ battery capacities, behavior, and segment access we present scenarios for California’s charging demand in 2030 with 8 million passenger electric vehicles. Peak charging demand ranged from 3.3 to 8.7 GW across scenarios. Furthermore, each was calculated in under 45 s on a laptop computer.

33 ADVANCED PROPULSION SYSTEMS↗

Predicting the viability of beta-lactamase: How folding and binding free energies correlate with beta-lactamase fitness

One of the long-standing holy grails of molecular evolution has been the ability to predict an organism’s fitness directly from its genotype. With such predictive abilities in hand, researchers would be able to more accurately forecast how organisms will evolve and how proteins with novel functions could be engineered, leading to revolutionary advances in medicine and biotechnology. In this work, we assemble the largest reported set of experimental TEM-1 β-lactamase folding free energies and use this data in conjunction with previously acquired fitness data and computational free energy predictions to determine how much of the fitness of β-lactamase can be directly predicted by thermodynamic folding and binding free energies. We focus upon β-lactamase because of its long history as a model enzyme and its central role in antibiotic resistance. Based upon a set of 21 β-lactamase single and double mutants expressly designed to influence protein folding, we first demonstrate that modeling software designed to compute folding free energies such as FoldX and PyRosetta can meaningfully, although not perfectly, predict the experimental folding free energies of single mutants. Interestingly, while these techniques also yield sensible double mutant free energies, we show that they do so for the wrong physical reasons. We then go on to assess how well both experimental and computational folding free energies explain single mutant fitness. We find that folding free energies account for, at most, 24% of the variance in β-lactamase fitness values according to linear models and, somewhat surprisingly, complementing folding free energies with computationally-predicted binding free energies of residues near the active site only increases the folding-only figure by a few percent. This strongly suggests that the majority of β-lactamase’s fitness is controlled by factors other than free energies. Overall, our results shed a bright light on to what extent the community is justified in using thermodynamic measures to infer protein fitness as well as how applicable modern computational techniques for predicting free energies will be to the large data sets of multiply-mutated proteins forthcoming.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Deep-Learning-Based Segmentation of Keyhole in In-Situ X-ray Imaging of Laser Powder Bed Fusion

In laser powder bed fusion processes, keyholes are the gaseous cavities formed where laser interacts with metal, and their morphologies play an important role in defect formation and the final product quality. The in-situ X-ray imaging technique can monitor the keyhole dynamics from the side and capture keyhole shapes in the X-ray image stream. Keyhole shapes in X-ray images are then often labeled by humans for analysis, which increasingly involves attempting to correlate keyhole shapes with defects using machine learning. However, such labeling is tedious, time-consuming, error-prone, and cannot be scaled to large data sets. To use keyhole shapes more readily as the input to machine learning methods, an automatic tool to identify keyhole regions is desirable. In this paper, a deep-learning-based computer vision tool that can automatically segment keyhole shapes out of X-ray images is presented. The pipeline contains a filtering method and an implementation of the BASNet deep learning model to semantically segment the keyhole morphologies out of X-ray images. The presented tool shows promising average accuracy of 91.24% for keyhole area, and 92.81% for boundary shape, for a range of test dataset conditions in Al6061 (and one AliSi10Mg) alloys, with 300 training images/labels and 100 testing images for each trial. Prospective users may apply the presently trained tool or a retrained version following the approach used here to automatically label keyhole shapes in large image sets.

36 MATERIALS SCIENCE↗

Comparison of Equilibrium and Nonequilibrium Approaches for Relative Binding Free Energy Predictions

Alchemical relative binding free energy calculations have recently found important applications in drug optimization. A series of congeneric compounds are generated from a preidentified lead compound, and their relative binding affinities to a protein are assessed in order to optimize candidate drugs. While methods based on equilibrium thermodynamics have been extensively studied, an approach based on nonequilibrium methods has recently been reported together with claims of its superiority. However, these claims pay insufficient attention to the basis and reliability of both methods. Here we report a comparative study of the two approaches across a large data set, comprising more than 500 ligand transformations spanning in excess of 300 ligands binding to a set of 14 diverse protein targets. Ensemble methods are essential to quantify the uncertainty in these calculations, not only for the reasons already established in the equilibrium approach but also to ensure that the nonequilibrium calculations reside within their domain of validity. If and only if ensemble methods are applied, we find that the nonequilibrium method can achieve accuracy and precision comparable to those of the equilibrium approach. Compared to the equilibrium method, the nonequilibrium approach can reduce computational costs but introduces higher computational complexity and longer wall clock times. There are, however, cases where the standard length of a nonequilibrium transition is not sufficient, necessitating a complete rerun of the entire set of transitions. This significantly increases the computational cost and proves to be highly inconvenient during large-scale applications. Our findings provide a key set of recommendations that should be adopted for the reliable implementation of nonequilibrium approaches to relative binding free energy calculations in ligand-protein systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Mass Dependence of Galaxy–Halo Alignment in LOWZ and CMASS

We measure the galaxy-ellipticity (GI) correlations for the Sloan Digital Sky Survey Data Release 12 LOWZ and CMASS samples with the shape measurements from the DESI Legacy Imaging Surveys. We model the GI correlations in an N-body simulation with our recent accurate stellar–halo mass relation from the Photometric object Around Cosmic webs (PAC) method. The large data set and our accurate modeling turns out an accurate measurement of the alignment angle between central galaxies and their host halos. We find that the alignment of central elliptical galaxies with their host halos increases monotonically with galaxy stellar mass or host halo mass, which can be well described by a power law for the massive galaxies. We also find that central elliptical galaxies are more aligned with their host halos in LOWZ than in CMASS, which might indicate an evolution of galaxy–halo alignment, though future studies are needed to verify this is not induced by the sample selections. In contrast, central disk galaxies are aligned with their host halos about 10 times more weakly in the GI correlation. These results have important implications for intrinsic alignment (IA) correction in weak lensing studies, IA cosmology, and theory of massive galaxy formation.

79 ASTRONOMY AND ASTROPHYSICS↗

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics ↗

Soil organic carbon is not just for soil scientists: measurement recommendations for diverse practitioners

Soil organic carbon (SOC) regulates terrestrial ecosystem functioning, provides diverse energy sources for soil microorganisms, governs soil structure, and regulates the availability of organically bound nutrients. Investigators in increasingly diverse disciplines recognize how quantifying SOC attributes can provide insight about ecological states and processes. Today, multiple research networks collect and provide SOC data, and robust, new technologies are available for managing, sharing, and analyzing large data sets. Here, we advocate that the scientific community capitalize on these developments to augment SOC data sets via standardized protocols. We describe why such efforts are important and the breadth of disciplines for which it will be helpful, and outline a tiered approach for standardized sampling of SOC and ancillary variables that ranges from simple to more complex. We target scientists ranging from those with little to no background in soil science to those with more soil-related expertise, and offer examples of the ways in which the resulting data can be organized, shared, and discoverable.

54 ENVIRONMENTAL SCIENCES↗

What's Left for a Computational Chemist To Do in the Age of Machine Learning?

Machine learning (ML) has become a central focus of the computational chemistry community. In this paper, I will first discuss my personal history in the field. Then I will provide a broader view of how this resurgence in ML interest echoes and advances upon earlier efforts. Although numerous changes have brought about this latest wave, one of the most significant is the increased accuracy and efficiency of low-cost methods (e.g., density functional theory or DFT) that have made it possible to generate large data sets for ML models. ML has also been used to bypass, guide, or improve DFT. The field of computational chemistry thus finds itself at a crossroads as ML both augments and supersedes traditional efforts. I will present what I believe the role of the computational chemist will be in this evolving landscape, with specific focus on my experience in the development of autonomous workflows in computational materials discovery for open-shell transition-metal chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A riming‐dependent parameterization of scattering by snowflakes using the self‐similar Rayleigh–Gans approximation

Abstract Riming is a key process of precipitation formation in ice‐containing clouds, but quantifying riming from observations is challenging, limiting our ability to evaluate the riming process in numerical weather models. One challenge for radar observations is that riming changes both the physical properties (mass, area cross‐section) and scattering properties of ice particles. These changes need to be implemented consistently as a function of riming in radar forward operators, which are required for retrievals and model evaluation in observation space. In this study, mass–size, cross‐section area–size, and backscattering cross‐section relations are developed as a function of the normalized rime mass for aggregates composed of various monomer types (columns, dendrites, needles, plates, and rosettes). The proposed framework allows us to simulate scattering properties of aggregated ice particles consistently as a function of riming in retrievals and radar forward operators. The parameterizations are developed from a large data set of simulated rimed aggregates of different sizes and monomer crystal types. The backscattering cross‐section parameterization (the “riming‐dependent parameterization”) is evaluated for radar frequencies of 35.6 and 94.0 GHz and is based on the Self‐Similar Rayleigh–Gans approximation (SSRGA), which is increasingly used to calculate microwave scattering of ice crystals and snowflakes. Compared with parameterizations from the literature that do not consider riming, the riming‐dependent parameterization leads to significantly smaller biases in terms of backscattering cross‐section. When using the particle masses and scattering properties of the individual particles simulated by the aggregation and riming model as a reference, the bias of our parameterization is below 1 dB when integrating over an exponential particle size distribution with sizes from 0.1–10 mm.

54 ENVIRONMENTAL SCIENCES↗

libwfa: Wavefunction analysis tools for excited and open‐shell electronic states

Abstract An open‐source software library for wavefunction analysis, libwfa, provides a comprehensive and flexible toolbox for post‐processing excited‐state calculations, featuring a hierarchy of interconnected visual and quantitative analysis methods. These tools afford compact graphical representations of various excited‐state processes, provide detailed insight into electronic structure, and are suitable for automated processing of large data sets. The analysis is based on reduced quantities, such as state and transition density matrices (DMs), and allows one to distill simple molecular orbital pictures of physical phenomena from intricate correlated wavefunctions. The implemented descriptors provide a rigorous link between many‐body wavefunctions and intuitive physical and chemical models, for example, exciton binding, double excitations, orbital relaxation, and polyradical character. A broad range of quantum‐chemical methods is interfaced with libwfa via a uniform interface layer in the form of DMs. This contribution reviews the structure of libwfa and highlights its capabilities by several representative use cases. This article is categorized under: Software > Quantum Chemistry Theoretical and Physical Chemistry > Spectroscopy

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NanoSIP: NanoSIMS Applications for Microbial Biology

High-resolution imaging with secondary ion mass spectrometry (nanoSIMS) has become a standard method in systems biology and environmental biogeochemistry and is broadly used to decipher ecophysiological traits of environmental microorganisms, metabolic processes in plant and animal tissues, and cross-kingdom symbioses. When combined with stable isotope-labeling—an approach we refer to as nanoSIP—nanoSIMS imaging offers a distinctive means to quantify net assimilation rates and stoichiometry of individual cell-sized particles in both low- and high-complexity environments. While the majority of nanoSIP studies in environmental and microbial biology have focused on nitrogen and carbon metabolism (using 15 N and 13 C tracers), multiple advances have pushed the capabilities of this approach in the past decade. The development of a high-brightness oxygen ion source has enabled high-resolution metal analyses that are easier to perform, allowing quantification of metal distribution in cells and environmental particles. New preparation methods, tools for automated data extraction from large data sets, and analytical approaches that push the limits of sensitivity and spatial resolution have allowed for more robust characterization of populations ranging from marine archaea to fungi and viruses. Further, NanoSIMS studies continue to be enhanced by correlation with orthogonal imaging and ‘omics approaches; when linked to molecular visualization methods, such as in situ hybridization and antibody labeling, these techniques enable in situ function to be linked to microbial identity and gene expression. Here we present an updated description of the primary materials, methods, and calculations used for nanoSIP, with an emphasis on recent advances in nanoSIMS applications, key methodological steps, and potential pitfalls.

07 ISOTOPE AND RADIATION SOURCES↗

Interactive Exploration of High-Dimensional Phase Diagrams

High-dimensional thermodynamic phase stability databases are becoming increasingly common due to the convergence of three recent trends: (i) the widespread interest in so-called “high-entropy” alloys, (ii) the availability of high-throughput computational assessments of phase stability in broad composition spaces and (iii) the ongoing development of ever-increasingly broad, multicomponent, multiphase CALPHAD databases. Although automated computational tools can readily process such high-dimensional data, scientists are often unable to visualize the relevant phase relations, an ability that is crucial to gaining an intuitive understanding of the stability constraints governing materials design. The present work addresses this need by providing algorithms that enable the interactive exploration of phase equilibria in high-dimensional spaces. These algorithms concentrate the complex nonlinear nonsmooth optimization needed into a preprocessing step that generates a large number of high-dimensional yet elementary graphical primitives. Furthermore, these primitives can then be cross-sectioned to yield 3-dimensional views in a computationally efficient manner that enables an interactive exploration of high-dimensional spaces. All of these operations are highly parallelizable, thus facilitating scaling of this method to large data sets.

36 MATERIALS SCIENCE↗

Piecewise linear approximation with minimum number of linear segments and minimum error: A fast approach to tighten and warm start the hierarchical mixed integer formulation

In several areas of economics and engineering, it is often necessary to fit discrete data points or approximate nonlinear functions with continuous functions. Piecewise linear (PWL) functions are a convenient way to achieve this. PWL functions can be modeled in mathematical problems using only linear and integer variables. Moreover, there is a computational benefit in using PWL functions that have the least possible number of segments. This work proposes a novel hierarchical mixed integer linear programming (MILP) formulation that identifies a continuous PWL approximation with minimum number of linear segments for a given target maximum error. The proposed MILP formulation also identifies the solution with the least maximum error among the solutions with minimum number of segments. Then, this work proposes a fast iterative algorithm that identifies non necessarily continuous PWL approximations by solving O(S log N) linear programming (LP) problems, where N is the number of data points and S is the minimum number of segments in the non necessarily continuous case. This work demonstrates that tight bounds for the MILP problem can be derived from these approximations. Next, a fast algorithm is introduced to transform a non necessarily continuous PWL approximation into a continuous one. Finally, the tight bounds and the continuous PWL approximations are used to tighten and warm start the MILP problem. The tightened formulation is shown in experimental results to be more efficient, especially for large data sets, with a solution time that is up to two orders of magnitude less than the existing literature.

97 MATHEMATICS AND COMPUTING↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Solving differential equations using deep neural networks

Recent work on solving partial differential equations (PDEs) with deep neural networks (DNNs) is presented. The paper reviews and extends some of these methods while carefully analyzing a fundamental feature in numerical PDEs and nonlinear analysis: irregular solutions. First, the Sod shock tube solution to the compressible Euler equations is discussed and analyzed. This analysis includes a comparison of a DNN-based approach with conventional finite element and finite volume methods, and demonstrates that the DNN is competitive in terms of degrees of freedom required for a given accuracy. Further, the DNN-based approach is extended to consider performance improvements and simultaneous parameter space exploration. Next, a shock solution to compressible magnetohydrodynamics (MHD) is solved for, and used in a scenario where experimental data is utilized to enhance a PDE system that is a priori insufficient to validate against the observed/experimental data. This is accomplished by enriching the model PDE system with source terms that are then inferred via supervised training with synthetic experimental data. The resulting DNN framework for PDEs enables straightforward system prototyping and natural integration of large data sets (be they synthetic or experimental), all while simultaneously enabling single-pass exploration of an entire parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗