Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “snapshot”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Inferring microbial co-occurrence networks from amplicon data: a systematic evaluation

Microbes commonly organize into communities consisting of hundreds of species involved in complex interactions with each other. 16S ribosomal RNA (16S rRNA) amplicon profiling provides snapshots that reveal the phylogenies and abundance profiles of these microbial communities. These snapshots, when collected from multiple samples, can reveal the co-occurrence of microbes, providing a glimpse into the network of associations in these communities. However, the inference of networks from 16S data involves numerous steps, each requiring specific tools and parameter choices. Moreover, the extent to which these steps affect the final network is still unclear. In this study, we perform a meticulous analysis of each step of a pipeline that can convert 16S sequencing data into a network of microbial associations. Through this process, we map how different choices of algorithms and parameters affect the co-occurrence network and identify the steps that contribute substantially to the variance. We further determine the tools and parameters that generate robust co-occurrence networks and develop consensus network algorithms based on benchmarks with mock and synthetic data sets. The Microbial Co-occurrence Network Explorer, or MiCoNE (available at https://github.com/segrelab/MiCoNE) follows these default tools and parameters and can help explore the outcome of these combinations of choices on the inferred networks. We envisage that this pipeline could be used for integrating multiple data sets and generating comparative analyses and consensus networks that can guide our understanding of microbial community assembly in different biomes.

16S rRNA↗

Vibronic and Environmental Effects in Simulations of Optical Spectroscopy

Including both environmental and vibronic effects is important for accurate simulation of optical spectra, but combining these effects remains computationally challenging. We outline two approaches that consider both the explicit atomistic environment and the vibronic transitions. Both phenomena are responsible for spectral shapes in linear spectroscopy and the electronic evolution measured in nonlinear spectroscopy. The first approach utilizes snapshots of chromophore-environment configurations for which chromophore normal modes are determined. We outline various approximations for this static approach that assumes harmonic potentials and ignores dynamic system-environment coupling. The second approach obtains excitation energies for a series of time-correlated snapshots. This dynamic approach relies on the accurate truncation of the cumulant expansion but treats the dynamics of the chromophore and the environment on equal footing. Both approaches show significant potential for making strides toward more accurate optical spectroscopy simulations of complex condensed phase systems.

Chemistry↗

Domain-decomposition nonlinear manifold reduced order model

This software combines nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD) techniques. NM-ROMs, which utilize a shallow, sparse autoencoder trained with full order model (FOM) snapshot data, approximate the FOM state on a nonlinear manifold. These models offer advantages over linear-subspace ROMs (LS-ROMs) particularly in scenarios with slowly decaying Kolmogorov n-width. However, the training of NM-ROMs involves a number of parameters that scale with the size of the FOM, and storing high-dimensional FOM snapshots can significantly increase the cost of ROM training for extreme-scale problems. To mitigate these costs, the software employs DD to partition the FOM into smaller subdomains, computes NM-ROMs for each, and then integrates these to form a global NM-ROM. This strategy offers multiple benefits: it enables parallel training of subdomain NM-ROMs, reduces the number of parameters needed, decreases the dimensional requirements of subdomain FOM training data, and allows for customization to the unique characteristics of each FOM subdomain. The use of a shallow, sparse autoencoder architecture in each subdomain NM-ROM facilitates the application of hyper-reduction (HR), simplifying the nonlinear complexities and enhancing computational speed. This software marks the inaugural application of NM-ROM combined with HR to a DD problem. It features an algebraic DD reformulation of the FOM, training of NM-ROMs with HR for each subdomain, and employs a sequential quadratic programming (SQP) solver for the evaluation of the coupled global NMROM. The effectiveness of the DD NM-ROM with HR is numerically demonstrated on the 2D steady-state Burgers' equation, showing an order of magnitude improvement in accuracy over the DD LS-ROM with HR.

Diaz, AlejandroN↗

PV Generation and Load Forecasting for Adjuntas PR Community Microgrids

Existing frameworks to forecast time-series photovoltaic (PV) output power and consumer load for microgrid operations and controls assume a near-continuous availability of real-time input features from the field assets such as PV inverters, energy meters, and weather station. These incoming data points are used to periodically retrain models and update forecast snapshots over a moving horizon window, be it one hour-ahead, one-day ahead, or one-week ahead. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. Hence, there is a need for programs that assume no availability of real-time microgrid asset data and still make reliable forecasts that can be used for decision-making. Such programs would be apt to function in extreme weather events such as hurricanes and would use lightweight recursive time-series models to independently forecast solar irradiance and ambient temperature, then compute PV power from those forecasts, as well as independently forecast consumer load. The codebase performs forecasting for the scenario of when the microgrid does not have a reliable access to forecasts or real-time observations of solar irradiance (I) and ambient temperature (AT) and load (Load) to be able to adequately forecast, in real-time, the PV power production or a business' load. In this case, using historical values of PV power and load, a univariate forecasting of generation and consumption are respectively made. The use-case in particular has two sub-scenarios: one, a normal 7-day ahead forecast where the unavailability of real-time data is assumed due to infrastructure issues such as loss of communication or sensor maintenance or service downtimes. Whereas a hurricane-caused unavailability of real-time data requires a second model trained specifically on historical hurricane days to be able to capture the extreme day behavior of generation in particular, and load if applicable. A gradient boosted regression tree comprises an ensemble of additive models that map between the input of historical values (be it irradiance, temperature, or load) and their corresponding output forecasts of a given horizon such that the individual learner predictions are summed up over the total number of such learners in the ensemble to produce an aggregate forecast. A weighting mechanism is applied to the training data in each iteration, where actual and forecast values are compared to penalize incorrect forecasts by increasing the weight and reducing it to reward correct forecasts. The code's benefits are that it: (a) accounts for a contingency where communication loss renders newly measured real-time data unavailable for model tuning and snapshot updates; (b) presents blind forecasting that recursively determines the next time-step value in a horizon using the forecast of the same attribute from a prior step; and (c) employs lightweight models that, once trained, can reliably generalize for different horizons, which make them suitable for enhancing the resilience of field microgrids prone to extreme events that encounter disruptions to data availability.

Sundararajan, Aditya [Oak Ridge National Laborator↗

A Nowcasting Approach for Low-Earth-Orbiting Hyperspectral Infrared Soundings within the Convective Environment

Low-Earth-orbiting (LEO) hyperspectral infrared (IR) sounders have significant yet untapped potential for characterizing thermodynamic environments of convective initiation and ongoing convection. While LEO soundings are of value to weather forecasters, the temporal resolution needed to resolve the rapidly evolving thermodynamics of the convective environment is limited. Here, we have developed a novel nowcasting methodology to extend snapshots of LEO soundings forward in time up to 6 h to create a product available within National Weather Service systems for user assessment. Our methodology is based on parcel forward-trajectory calculations from the satellite-observing time to generate future soundings of temperature (T) and specific humidity (q) at regularly gridded intervals in space and time. The soundings are based on NOAA-Unique Combined Atmospheric Processing System (NUCAPS) retrievals from the Suomi National Polar-Orbiting Partnership (Suomi NPP) and NOAA-20 satellite platforms. The tendencies of derived convective available potential energy (CAPE) and convective inhibition (CIN) are evaluated against gridded, hourly accumulated rainfall obtained from the Multi-Radar Multi-Sensor (MRMS) observations for 24 hand-selected cases over the contiguous United States. Areas with forecast increases in CAPE (reduced CIN) are shown to be associated with areas of precipitation. The increases in CAPE and decreases in CIN are largest for areas that have the heaviest precipitation and are statistically significant compared to areas without precipitation. These results imply that adiabatic parcel advection of LEO satellite sounding snapshots forward in time are capable of identifying convective initiation over an expanded temporal scale compared to soundings used only during the LEO satellite overpass time.

54 ENVIRONMENTAL SCIENCES↗

Pseudonymized User-Perspective Summit Login Node Data for 2020 and 2021

This dataset contains hourly snapshot data from each of the 5 login nodes of the Summit supercomputer at Oak Ridge Leadership Computing Facility (OLCF) over a period of 2 years, starting January 2020 and ending after December 2021. The snapshots include lists of currently logged-in users, CPU and memory usage, status of users' batch jobs, and disk usage statistics. Usernames, project identifiers, and file paths have been pseudonymized in order to allow studies of user behavior without divulging Personally Identifiable Information (PII).

97 MATHEMATICS AND COMPUTING↗

SST-TG-P1F4R3200: Decaying Stably-Stratified Turbulence (SST), Initialized Using Taylor-Green Vortices (TG) at Prandtl Number Pr=1, Froude Number Fr=4, Reynolds Number Re=3200

This dataset comprises direct numerical simulations (DNS) of decaying stably-stratified turbulence influenced by a linear background density gradient, initialized using an array of Taylor-Green vortices, as described in [Riley & de Bruyn Kops (2003)](https://doi.org/10.1063/1.1578077). The initial Prandtl, Froude, and Reynolds numbers are (Pr, Fr, Re) = (1, 4, 3200). A total of 15,000 snapshots are recorded at uniform time intervals, each with a spatial resolution of 512x512x256 grid points. Four flow variables are associated with each snapshot: the three velocity components (u,v,w) and the perturbed density field (rho) away from the background gradient. All fields are stored in binary format (32-bit little-endian), each with a size of 255 MB, yielding a total dataset size of 15.3 TB. Further details are referenced in the attached README file, and a current list of publications and associated analysis tools are provided at https://stratified-turbulence.github.io/web/.

42 ENGINEERING↗

SST-TG-P50F4R3200: Decaying Stably-Stratified Turbulence (SST), Initialized Using Taylor-Green Vortices (TG) at Prandtl Number Pr=50, Froude Number Fr=4, Reynolds Number Re=3200

This dataset comprises direct numerical simulations (DNS) of decaying stably-stratified turbulence influenced by a linear background density gradient, initialized using an array of Taylor-Green vortices, extending the Pr=1 simulations performed in [Riley and de Bruyn Kops (2003)](https://doi.org/10.1063/1.1578077). The initial Prandtl, Froude, and Reynolds numbers are (Pr, Fr, Re) = (50, 4, 3200). A total of 1,680 snapshots are recorded at uniform time intervals, each with a spatial resolution of 3584x3584x1792 grid points. Four flow variables are associated with each snapshot: the three velocity components (u,v,w) and the perturbed density field (rho) away from the background gradient. All fields are stored in binary format (32-bit little-endian), each with a size of 85.8 GB, yielding a total dataset size of 577 TB. Further details are referenced in the attached README file, and a current list of publications and associated analysis tools are provided at https://stratified-turbulence.github.io/web/.

42 ENGINEERING↗

SST-TG-P7F4R3200: Decaying Stably-Stratified Turbulence (SST), Initialized Using Taylor-Green Vortices (TG) at Prandtl Number Pr=7, Froude Number Fr=4, Reynolds Number Re=3200

This dataset comprises direct numerical simulations (DNS) of decaying stably-stratified turbulence influenced by a linear background density gradient, initialized using an array of Taylor-Green vortices, extending the Pr=1 simulations performed in [Riley and de Bruyn Kops (2003)](https://doi.org/10.1063/1.1578077). The initial Prandtl, Froude, and Reynolds numbers are (Pr, Fr, Re) = (7, 4, 3200). A total of 15,250 snapshots are recorded at uniform time intervals, each with a spatial resolution of 1280x1280x640 grid points. Four flow variables are associated with each snapshot: the three velocity components (u,v,w) and the perturbed density field (rho) away from the background gradient. All fields are stored in binary format (32-bit little-endian), each with a size of 4 GB, yielding a total dataset size of 244 TB. Further details are referenced in the attached README file, and a current list of publications and associated analysis tools are provided at https://stratified-turbulence.github.io/web/.

42 ENGINEERING↗

Cholla Galactic OutfLow Simulations (CGOLS)

These datasets contain full hydro-field snapshots from the galactic outflow simulations in the CGOLS suite, models I-V. The datasets were generated using the Cholla hydrodynamics code (https://github.com/cholla-hydro/cholla); descriptions of the models are in the associated publications (Schneider & Robertson 2018, ApJ; Schneider et al. 2018, ApJ; Schneider et al. 2020, ApJ; and Schneider & Mao, 2024, ApJ). Each hdf5 dataset is numbered according to the simulation time of the snapshot, in Myr. Fields include density, x momentum, y momentum, z momentum, total energy, and thermal energy (for models I - III), as well as a passive scalar field (models IV and V). 2 dimensional density and temperature projections, as well as slices along each midplane are also included if they exist.

79 ASTRONOMY AND ASTROPHYSICS↗

Single-molecule 3D imaging of HIV cellular entry by liquid-phase electron tomography

Enveloped viruses, including human immunodeficiency virus (HIV) and SARS-CoV-2, target cells through membrane fusion process. The detailed understanding of the process is sought after for vaccine development but remains elusive due to current technique limitations for direct three-dimensional (3D) imaging of an individual virus during its viral entry. Recently, we developed a simple specimen preparation method for real-time imaging of metal dynamic liquid-vaper interface at nanometer resolution by transmission electron microscopy (TEM). Here, we extended this method to study biology sample through snapshot 3D structure of a single HIV (pseudo-typed with the envelope glycoprotein of vesicular stomatitis virus, VSV-G) at its intermediate stage of viral entry to HeLa cells in a liquid-phase environment. By individual-particle electron tomography (IPET), we found the viral surface release excess lipids with unbound viral spike proteins forming ~50-nm nanoparticles instead of merging cell membrane. Moreover, the spherical-shape shell formed by matrix proteins underneath the viral envelope does not disassemble into a cone shape right after fusion. Further, the snapshot 3D imaging of a single virus provides us a direct structure-based understanding of the viral entry mechanism, which can be used to examine other viruses to support the development of vaccines combatting the current ongoing pandemic.

Kong, Lingli↗

Implementing Software Resiliency in HPX for Extreme Scale Computing

The DOE Office of Science Exascale Computing Project (ECP) outlines the next milestones in the supercomputing domain. The target computing systems under the project will deliver 10x performance while keeping the power budget under 30 megawatts. With such large machines, the need to make applications resilient has become paramount. The benefits of adding resiliency to mission critical and scientific applications, includes the reduced cost of restarting the failed simulation both in terms of time and power. Most of the current implementation of resiliency at the software level makes use of a Coordinated Checkpoint and Restart (C/R). This technique of resiliency generates a consistent global snapshot, also called a checkpoint. Generating snapshots involves global communication and coordination and is achieved by synchronizing all running processes. The generated checkpoint is then stored in some form of persistent storage. On failure detection, the runtime initiates a global rollback to the most recent previously saved checkpoint. This involves aborting all running processes, rolling them back to the previous state and restarting them.

97 MATHEMATICS AND COMPUTING↗

Identification of Faults Susceptible to Induced Seismicity (Final Report)

Central to the work documented in this report is the capability of geocellular models to represent the geologic conceptual model updated with fault identification from machine learning and joint inversion modeling of microseismic data measured and recorded as a consequence of CO 2 injection at a field demonstration site: the Illinois Basin - Decatur Project (IBDP). This work required seven unique geocellular models with 100s of simulated variations to gain a very high degree of confidence in the identification of geologic features present that contributed to induced microseismicity at IBDP. All forward modeling: pressure modeling, stress modeling, and seismic modeling used the same geologic conceptual model and representations of that model at different scales. The pressure modeling and poroelastic modeling created “snapshots” of pore pressure and stress field changes at different times during CO 2 injection, in which microseismic events were clustered (in time). These pressure and stress snapshots, within the framework and architecture of the geologic conceptual model via the geocellular model, informed the single fault and fault network models to ascertain the likelihood of fault movement (seismic or aseismic). The outcomes of the pressure, stress, and fault/fault network (seismic) modeling confirmed that the faults in the geologic conceptual model in Task 2 were likely the source of microseismic events measured at IBDP and acted as conduits for pressure to be transmitted from the injection interval into the Precambrian crystalline basement rock. This closely coordinated and integrated unique modeling approach was conducted to prove the viability of our proposed workflow 1) to better resolve crystalline basement faults, 2) detect subseismic faults that could be activated by injection, 3) increase the certainty in fault detection and their susceptibility to release seismic energy, and 4) understand transmission of pressure vertically from the well to the underlying fractured crystalline basement. The proposed methodology was effective in guiding an iterative process of calibrating forward modeling results based on similar geocellular models while honoring the geologic conceptual model (i.e., characterization data and knowledge of regional geology); this led to higher level of certainty in the identification of fault/faults zones to control seismicity and transmission of pressure to the regions of recorded and located injection induced seismicity.

58 GEOSCIENCES↗

Augmented reduced order models for turbulence

The authors introduce an augmented-basis method (ABM) to stabilize reduced-order models (ROMs) of turbulent incompressible flows. The method begins with standard basis functions derived from proper orthogonal decomposition (POD) of snapshot sets taken from a full-order model. These are then augmented with divergence-free projections of a subset of the nonlinear interaction terms that constitute a significant fraction of the time-derivative of the solution. The augmenting bases, which are rich in localized high wavenumber content, are better able to dissipate turbulent kinetic energy than the standard POD bases. Several examples illustrate that the ABM significantly out-performs L 2 -, H 1 - and Leray-stabilized POD ROM approaches. The ABM yields accuracy that is comparable to constraint-based stabilization approaches yet is suitable for parametric model-order reduction in which one uses the ROM to evaluate quantities of interests at parameter values that differ from those used to generate the full-order model snapshots. Several numerical experiments point to the importance of localized high wavenumber content in the generation of stable, accurate, and efficient ROMs for turbulent flows.

Kaneko, Kento↗

Can strong substorm-associated MeV electron injections be an important cause of large radiation belt enhancements?

It has become well-established that strong outer radiation belt enhancements are due to wave-driven electron energization by whistler-mode chorus waves. However, in this study, we examine strong MeV electron injections on 10 July 2019 and find substantial evidence that such injections may be a crucial contributor to outer radiation belt enhancement events. For such an examination, it is essential to precisely separate temporal flux changes from spatial variations observed as Van Allen Probes move along their orbits. Employing a new “hourly snapshot” analysis approach, we discover unprecedented details of electron flux evolutions that suggest that for this event, the outer belt enhancement was not continuous but instead intermittent, mostly composed of 4 large discrete injection-driven flux increases. The injections appear as sharp flux increases when observed near apogee. Otherwise, by comparing hourly snapshots for different times, we infer injections and infer temporally stable fluxes between injections, despite strong and continuous chorus emission. The fast and intermittent electron flux growth successively extending earthwards implies cumulative outer belt enhancement via a series of repetitive inward transport associated with injection-induced electric fields.

79 ASTRONOMY AND ASTROPHYSICS↗

The Evolution of Volatile Memory Forensics

The collection and analysis of volatile memory is a vibrant area of research in the cybersecurity community. The ever-evolving and growing threat landscape is trending towards fileless malware, which avoids traditional detection but can be found by examining a system’s random access memory (RAM). Additionally, volatile memory analysis offers great insight into other malicious vectors. It contains fragments of encrypted files’ contents, as well as lists of running processes, imported modules, and network connections, all of which are difficult or impossible to extract from the file system. For these compelling reasons, recent research efforts have focused on the collection of memory snapshots and methods to analyze them for the presence of malware. However, to the best of our knowledge, no current reviews or surveys exist that systematize the research on both memory acquisition and analysis. We fill that gap with this novel survey by exploring the state-of-the-art tools and techniques for volatile memory acquisition and analysis for malware identification. For memory acquisition methods, we explore the trade-offs many techniques make between snapshot quality, performance overhead, and security. For memory analysis, we examined the traditional forensic methods used, including signature-based methods, dynamic methods performed in a sandbox environment, as well as machine learning-based approaches. We summarize the currently available tools, and suggest areas for more research.

Nyholm, Hannah↗

Hierarchical Fractional Advection-Dispersion Equation (FADE) to Quantify Anomalous Transport in River Corridor over a Broad Spectrum of Scales: Theory and Applications

Fractional calculus-based differential equations were found by previous studies to be promising tools in simulating local-scale anomalous diffusion for pollutants transport in natural geological media (geomedia), but efficient models are still needed for simulating anomalous transport over a broad spectrum of scales. This study proposed a hierarchical framework of fractional advection-dispersion equations (FADEs) for modeling pollutants moving in the river corridor at a full spectrum of scales. Applications showed that the fixed-index FADE could model bed sediment and manganese transport in streams at the geomorphologic unit scale, whereas the variable-index FADE well fitted bedload snapshots at the reach scale with spatially varying indices. Further analyses revealed that the selection of the FADEs depended on the scale, type of the geomedium (i.e., riverbed, aquifer, or soil), and the type of available observation dataset (i.e., the tracer snapshot or breakthrough curve (BTC)). When the pollutant BTC was used, a single-index FADE with scale-dependent parameters could fit the data by upscaling anomalous transport without mapping the sub-grid, intermediate multi-index anomalous diffusion. Pollutant transport in geomedia, therefore, may exhibit complex anomalous scaling in space (and/or time), and the identification of the FADE’s index for the reach-scale anomalous transport, which links the geomorphologic unit and watershed scales, is the core for reliable applications of fractional calculus in hydrology.

97 MATHEMATICS AND COMPUTING↗

Characterizing Families of Spectral Similarity Scores and Their Use Cases for Gas Chromatography–Mass Spectrometry Small Molecule Identification

Metabolomics provides a unique snapshot into the world of small molecules and the complex biological processes that govern the human, animal, plant, and environmental ecosystems encapsulated by the One Health modeling framework. However, this “molecular snapshot” is only as informative as the number of metabolites confidently identified within it. The spectral similarity (SS) score is traditionally used to identify compound(s) in mass spectrometry approaches to metabolomics, where spectra are matched to reference libraries of candidate spectra. Unfortunately, there is little consensus on which of the dozens of available SS metrics should be used. This lack of standard SS score creates analytic uncertainty and potentially leads to issues in reproducibility, especially as these data are integrated across other domains. In this work, we use metabolomic spectral similarity as a case study to showcase the challenges in consistency within just one piece of the One Health framework that must be addressed to enable data science approaches for One Health problems. Here, using a large cohort of datasets comprising both standard and complex datasets with expert-verified truth annotations, we evaluated the effectiveness of 66 similarity metrics to delineate between correct matches (true positives) and incorrect matches (true negatives). We additionally characterize the families of these metrics to make informed recommendations for their use. Our results indicate that specific families of metrics (the Inner Product, Correlative, and Intersection families of scores) tend to perform better than others, with no single similarity metric performing optimally for all queried spectra. This work and its findings provide an empirically-based resource for researchers to use in their selection of similarity metrics for GC-MS identification, increasing scientific reproducibility through taking steps towards standardizing identification workflows.

59 BASIC BIOLOGICAL SCIENCES↗