Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine Learning for Distributed Acoustic Sensing data (MLDAS) v1.0.1

MLDAS is a Python-written package for exploratory data analysis and deep learning training on Distributed Acoustic Sensing data. The machine learning tools are powered by the PyTorch library and designed to work efficiently on large scale datasets using parallel computing. Various SLURM scripts as well as a tutorial have also been made available to allow geophysicists to quickly and easily implement the available tools in their analysis workflow on supercomputer facilities.

Dumont, Vincent↗

Massively Parallel and Portable Genomic Sequence Analysis

Massively Parallel and Portable Genomic Sequence Analysis (mappgene) is a sequencing analysis workflow for high performance computing. It incorporates novel technologies to simplify and accelerate genetics research.

Moon, Josephy↗

NOODLES [SWR-22-78]

NOODLES is a protocol specification for collaborative visualization. It allows software tools of any type to participate in an analysis or visualization session. Use cases include, but are not limited to, distributed analysis, workflow interoperability, computational steering, etc.

Brunhart-Lupo, Nicholas↗

ESS-DIVE Reporting Format for Amplicon Abundance Table

While standardized sequencing data is available in public repositories and efforts such as MIxS for common sample collection and processing metadata are well established, the lack of common bioinformatic processing metadata has hindered the ability to do large-scale metaanalyses and the potential for data re-use by non-experts such as ecosystem, watershed, or earth system modelers. To address this need for Department of Energy researchers, we have developed an amplicon reporting format which captures both sample preparation and bioinformatic processing metadata and stores processed amplicon data as a paired abundance table and sequencing file to maximize the potential for re-use of these data. To aid in the adoption of accessible and reproducible analysis workflows, this reporting format was developed in concert with amplicon functionality within the Department of Energy’s Systems Biology Knowledgebase (KBase) to ensure common data and metadata requirements and facilitate seamless transfer between these platforms.This dataset contains support documentation for the amplicon reporting format (README.md and instructions.md), templates for both bioinformatic and sequencing metadata (amplicon_bioinformatic_metadata_template_2021_10_03.csv and amplicon_sequencing_metadata_template_2021_10_03.csv), a crosswalk indicating how this reporting format relates to the current MIxS format (ESSDIVE-MIxS_crosswalk.csv), a list of available instrument terms (amplicon_seq_instrument_terms_2021_10_03.csv), a map between QIIME2 parameter settings and metadata fields (amplicon_qiime2_plugin_metadata_map.csv), a data dictionary (amplicon_CSV_dd.csv), and file-level metadata (amplicon_FLMD.csv).

54 ENVIRONMENTAL SCIENCES↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Evaluating Symmetry in 3-dimensional Simulation Data

The increasing use of three-dimensional simulations has led to an increased need for effective verification and validation, or V&V. Unfortunately, the high computational cost associated with three-dimensional simulations makes many standard V&V techniques cost-prohibitive. Given these circumstances, we must pay close attention to V&V metrics that can be evaluated using only a single simulation. One of the simplest of these metrics is the evaluation of physical symmetries. This report details a method of computing the deviation from symmetry in a three-dimensional dataset with an arbitrarily distributed mesh. Examples are presented for spherical and axisymmetric systems using the Paraview post-processing analysis workflow.

97 MATHEMATICS AND COMPUTING↗

Augmented Human Analysis (AHA)

Radio frequency (RF) signal monitoring generally emphasizes intentionally generated signals, such as WiFi, Bluetooth, or cellular transmissions. However, electronic devices also produce unintended radiated emissions (UREs), which could also be useful in RF spectrum analysis. In either case, deriving intelligence from RF signals is typically a human-intensive process requiring significant domain knowledge. In the Augmented Human Analysis (AHA) project, we investigate the utility of dimensionally aligned signal projection (DASP) and machine learning (ML) algorithms for accelerating RF analysis workflows. We find that while DASP algorithms can indeed highlight signal characteristics relevant for classification tasks, the choice of algorithmic hyperparameters greatly affects performance. To address this challenge, we evaluate the quality of DASP outputs using the silhouette score, which measures how well data points cluster; high silhouette scores indicate good clustering, and thus good hyperparameter values. This approach is critical for machine learning pipelines as the DASP parameters cannot be directly optimized during model training. By identifying good DASP parameters, and thus good DASP outputs, as a preprocessing step, we can decrease the amount of effort required for downstream ML model training. We demonstrate our workflow using a dataset of UREs from common household devices, showing that even without the aid of ML, proper selection of DASP parameters enables clustering by device type.

42 ENGINEERING↗

Powered By ERAD [Slides]

Energy Resilience Analysis for Distribution Power System (ERAD) is a free, open-source Python toolkit for estimating the energy and service impacts of hazards like earthquakes and flooding. It uses a graph-based approach to capture high resolution connectivity among the grid, critical services, and customers and rapidly compute household level metrics and aggregated statistics across large distribution systems. It uses asset fragility curves that relate hazard severity to survival probability for power system equipment including cables, transformers, substations, etc. The tool is designed to be modular and extensible, allowing it to interface with third-party hazard simulators and integrate into broader resilience analysis workflows. ERAD enables researchers, students, communities, distribution utilities, and other stakeholders to understand hazard impacts and evaluate the effectiveness of different programs to improve energy resilience. The webinar was hosted by NLR researcher Aadil Latif.

24 POWER TRANSMISSION AND DISTRIBUTION↗

DEIMoS: An Open-Source Tool for Processing High-Dimensional Mass Spectrometry Data

We present DEIMoS: Data Extraction for Integrated Multidimensional Spectrometry, a Python application programming interface and command-line tool for high-dimensional mass spectrometry data analysis workflows, offering ease of development and access to efficient algorithmic implementations. Functionality includes feature detection, feature alignment, collision cross section calibration, isotope detection, and MS/MS spectral deconvolution, with the output comprising detected features aligned across study samples and characterized by mass, CCS, tandem mass spectra, and isotopic signature. Notably, DEIMoS operates on N-dimensional data, largely agnostic to acquisition instrumentation: algorithm implementations utilize all dimensions simultaneously to (i) offer greater separation between features, improving detection sensitivity, (ii) increase alignment/feature matching confidence among datasets, and (iii) mitigate convolution artifacts in tandem mass spectra. We demonstrate DEIMoS with LC-IMS-MS/MS data, demonstrating the advantages of a multidimensional approach in each data processing step.

Colby, Sean M↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

Wavelet Analysis of GPR Data for Belowground Mass Assessment of Sorghum Hybrid for Soil Carbon Sequestration

Among many agricultural practices proposed to cut carbon emissions in the next 30 years is the deposition of carbon in soils as plant matter. Adding rooting traits as part of a sequestration strategy would result in significantly increased carbon sequestration. Integrating these traits into production agriculture requires a belowground phenotyping method compatible with high-throughput breeding (i.e., rapid, inexpensive, reliable, and non-destructive). However, methods that fulfill these criteria currently do not exist. We hypothesized that ground-penetrating radar (GPR) could fill this need as a phenotypic selection tool. In this study, we employed a prototype GPR antenna array to scan and discriminate the root and rhizome mass of the perennial sorghum hybrid PSH09TX15. B-scan level time/discrete frequency analyses using continuous wavelet transform were utilized to extract features of interest that could be correlated to the biomass of the subsurface roots and rhizome. Time frequency analysis yielded strong correlations between radar features and belowground biomass (max R −0.91 for roots and −0.78 rhizomes, respectively) These results demonstrate that continued refinement of GPR data analysis workflows should yield an applicable phenotyping tool for breeding efforts in contexts where selection is otherwise impractical.

Wolfe, Matthew↗

Distinguishing Provenance Equivalence of Earth Science Data

Reproducibility of scientific research relies on accurate and precise citation of data and the provenance of that data. Earth science data are often the result of applying complex data transformation and analysis workflows to vast quantities of data. Provenance information of data processing is used for a variety of purposes, including understanding the process and auditing as well as reproducibility. Certain provenance information is essential for producing scientifically equivalent data. Capturing and representing that provenance information and assigning identifiers suitable for precisely distinguishing data granules and datasets is needed for accurate comparisons. This paper discusses scientific equivalence and essential provenance for scientific reproducibility. We use the example of an operational earth science data processing system to illustrate the application of the technique of cascading digital signatures or hash chains to precisely identify sets of granules and as provenance equivalence identifiers to distinguish data made in an an equivalent manner.

Tilmes, Curt↗

Unstructured Grid Adaptation and Solver Technology for Turbulent Flows

Unstructured grid adaptation is a tool to control Computational Fluid Dynamics (CFD) discretization error. However, adaptive grid techniques have made limited impact on production analysis workflows where the control of discretization error is critical to obtaining reliable simulation results. Issues that prevent the use of adaptive grid methods are identified by applying unstructured grid adaptation methods to a series of benchmark cases. Once identified, these challenges to existing adaptive workflows can be addressed. Unstructured grid adaptation is evaluated for test cases described on the Turbulence Modeling Resource (TMR) web site, which documents uniform grid refinement of multiple schemes. The cases are turbulent flow over a Hemisphere Cylinder and an ONERA M6Wing. Adaptive grid force and moment trajectories are shown for three integrated grid adaptation processes with Mach interpolation control and output error based metrics. The integrated grid adaptation process with a finite element (FE) discretization produced results consistent with uniform grid refinement of fixed grids. The integrated grid adaptation processes with finite volume schemes were slower to converge to the reference solution than the FE method. Metric conformity is documented on grid/metric snapshots for five grid adaptation mechanics implementations. These tools produce anisotropic boundary conforming grids requested by the adaptation process.

Park, Michael A.↗

Verification of Unstructured Grid Adaptation Components

Adaptive unstructured grid techniques have made limited impact on production analysis workflows where the control of discretization error is critical to obtaining reliable simulation results. Recent progress has matured a number of independent implementations of flow solvers, error estimation methods, and anisotropic grid adaptation mechanics. Known differences and previously unknown differences in grid adaptation components and their integrated processes are identified here for study. Unstructured grid adaptation tools are verified using analytic functions and the Code Comparison Principle. Three analytic functions with different smoothness properties are adapted to show the impact of smoothness on implementation differences. A scalar advection-diffusion problem with an analytic solution that models a boundary layer is adapted to test individual grid adaptation components. Laminar flow over a delta wing and turbulent flow over an ONERA M6 wing are verified with multiple, independent grid adaptation procedures to show consistent convergence to fine-grid forces and a moment. The scalar problems illustrate known differences in a grid adaptation component implementation and a previously unknown interaction between components. The wing adaptation cases in the current study document a clear improvement to existing grid adaptation procedures. The stage is set for the infusion of verified grid adaptation into production fluid flow simulations.

Park, Michael A.↗

Verification of Unstructured Grid Adaptation Components

Adaptive unstructured grid techniques have made limited impact on production analysis workflows where the control of discretization error is critical to obtaining reliable simulation results. Recent progress has matured a number of independent implementations of flow solvers, error estimation methods, and anisotropic grid adaptation mechanics. Known differences and previously unknown differences in grid adaptation components and their integrated processes are identified here for study. Unstructured grid adaptation tools are verified using analytic functions and the Code Comparison Principle. Three analytic functions with different smoothness properties are adapted to show the impact of smoothness on implementation differences. A scalar advection-diffusion problem with an analytic solution that models a boundary layer is adapted to test individual grid adaptation components. The scalar problems illustrate known differences in a grid adaptation component implementation and a previously unknown interaction between components. Laminar flow over a delta wing is verified with multiple, independent grid adaptation procedures to show consistent convergence to fine-grid forces and pitching moment.

Park, Michael A.↗

Verification of Viscous Goal-Based Anisotropic Mesh Adaptation

Adaptive unstructured mesh techniques have a limited, but growing impact on production analysis workflows where the control of discretization error is critical to obtaining reliable simulation results. Recent progress has matured a number of independent implementations of flow solvers, error estimation methods, and anisotropic mesh adaptation mechanics. Anisotropic metric construction methods are evaluated with analytically defined primal and adjoint fields. This allows the comparison of different metric formulations and different implementations of the same formulation without the complications of a flow and adjoint solution method. Unstructured mesh adaptation tools are verified by comparison on analytic primal and dual field before verification on benchmark aerodynamics cases. The documentation of these verification exercises helps to prepare these goal-based methods for routine use in production simulation workflows.

Mesh adaptation↗

Maximizing Spaceflight Biological Data with Omics Analytics: The NASA GeneLab Database

NASA’s GeneLab includes an open-access repository of some 250+ omics datasets generated by biological experiments relevant to spaceflight including simulated cosmic radiation and microgravity. In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics background, GeneLab has become a knowledgebase platform converting raw genetic and proteomic signatures found in flight samples into biological and physiological meanings. A large community of more than 100 scientists has rallied behind GeneLab and organized into four Analysis Working Groups (AWGs: Animal, Plant, Microbe, and Multi-Omics). Together, the AWGs have gained scientific recognition worldwide by establishing a consortium in charge of adopting new complex standards for data analysis workflows and omics sample processing in a rapidly evolving field. We will demonstrate the usage of the repository with smart search capability, an online controlled-access toolshed "Galaxy" to process user data with vetted standard workflows, a workspace for data sharing and a data submission portal with ontology control for better metadata curation. The GeneLab visualization portal will also be demonstrated, showing how anyone without formal training in bioinformatics can now browse the space biology omics data to discover new biology and potential solutions to improve life in space.

Sylvain Vincent Costes↗

GeneLab: The NASA System Biology Platform for Space Omics Repository, Analysis and Visualization

NASA’s GeneLab includes an open-access repository of some 250+ omics datasets generated by biological experiments relevant to spaceflight including simulated cosmic radiation and microgravity. In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics background, GeneLab has become a knowledgebase platform converting raw genetic and proteomic signatures found in flight samples into biological and physiological meanings. A large community of more than 100 scientists has rallied behind GeneLab and organized into four Analysis Working Groups (AWGs: Animal, Plant, Microbe, and Multi-Omics). Together, the AWGs have gained scientific recognition worldwide by establishing a consortium in charge of adopting new complex standards for data analysis workflows and omics sample processing in a rapidly evolving field. We will demonstrate the usage of the repository with smart search capability, an online controlled-access toolshed "Galaxy" to process user data with vetted standard workflows, a workspace for data sharing and a data submission portal with ontology control for better metadata curation. The GeneLab visualization portal will also be demonstrated, showing how anyone without formal training in bioinformatics can now browse the space biology omics data to discover new biology and potential solutions to improve life in space.

GeneLab↗