Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

A Processing and Analytics System for Microscopy Data Workflows: The Pycroscopy Ecosystem of Packages

Major advancements in fields as diverse as biology and quantum computing have relied on a multitude of microscopy techniques. Despite the considerable proliferation of these instruments, significant bottlenecks remain in terms of processing, analysis, storage, and retrieval of the acquired datasets. Aside from lack of file standards, individual domain-specific analysis packages are often disjoint from the underlying datasets, and thus keeping track of analysis and processing steps remains tedious for the end-user, hampering reproducibility. Here, in this study, the pycroscopy ecosystem of packages is introduced, an open-source python-based ecosystem underpinned by a common data model. The data model, termed the N-dimensional spectral imaging data format, is realized in pycroscopy's sidpy package. This package is built on top of dask arrays, thus leveraging dask array attributes, but expanding them to accelerate microscopy relevant analysis and visualization. Several examples of the use of the pycroscopy ecosystem to create workflows for data ingestion and analysis of scanning transmission electron microscopy (STEM) and scanning probe microscopy data are shown. Adoption of such standardized routines will be critical to usher in the next generation of autonomous instruments where processing, computation, and meta-data storage will be critical to overall experimental operations.

97 MATHEMATICS AND COMPUTING↗

An efficient cre‐based workflow for genomic integration and expression of large biosynthetic pathways in Eubacterium limosum

Abstract Acetogenic Clostridia are obligate anaerobes that have emerged as promising microbes for the renewable production of biochemicals owing to their ability to efficiently metabolize sustainable single‐carbon feedstocks. Additionally, Clostridia are increasingly recognized for their biosynthetic potential, with recent discoveries of diverse secondary metabolites ranging from antibiotics to pigments to modulators of the human gut microbiota. Lack of efficient methods for genomic integration and expression of large heterologous DNA constructs remains a major challenge in studying biosynthesis in Clostridia and using them for metabolic engineering applications. To overcome this problem, we harnessed chassis‐independent recombinase‐assisted genome engineering (CRAGE) to develop a workflow for facile integration of large gene clusters (>10 kb) into the human gut acetogen Eubacterium limosum . We then integrated a non‐ribosomal peptide synthetase gene cluster from the gut anaerobe Clostridium leptum , which previously produced no detectable product in traditional heterologous hosts. Chromosomal expression in E. limosum without further optimization led to production of phevalin at 2.4 mg/L. These results further expand the molecular toolkit for a highly tractable member of the Clostridia, paving the way for sophisticated pathway engineering efforts, and highlighting the potential of E. limosum as a Clostridial chassis for exploration of anaerobic natural product biosynthesis.

Sanford, Patrick A.↗

A Tip-based Workflow for Sensitive IMAC-based Low Nanogram Level Phosphoproteomics

Analyzing the phosphoproteome at nanoscale poses a significant challenge, mainly due to the substantial sample loss from non-specific surface adsorption during the enrichment of low stoichiometric phosphopeptides. Here, we describe a tandem tip-based phosphoproteomics sample preparation method capable of sequential sample cleanup and enrichment without the need for additional sample transfer, thereby minimizing sample loss. Integration of this method to our recently developed SOP (Surfactant-assisted One-Pot sample preparation) and iBASIL (improved Boosting to Amplify Signal with Isobaric Labeling) approaches creates a streamlined workflow, enabling sensitive, high-throughput nanoscale phosphoproteomics measurements.

Phosphoproteome, Immobilized metal ion affinity ch↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗

A workflow for automatic generation and efficient refinement of individual pressure-dependent networks

Manual chemical kinetic model construction often requires modelists to at least implicitly guess all possible decay pathways and their relative fluxes for all important species. We present a workflow and associated tool for automatic generation and re finement of pressure-dependent networks that should enable modelists to efficiently and comprehensively identify decay pathways and estimate respective parameters for chemical species. This tool combines the capabilities of the Reaction Mechanism Generator (RMG) software to generate possible reaction paths, determine thermochemistry, approximate rate coefficients and estimate frequencies with the capabilities of the Arkane software to use quantum chemical parameters to compute thermochemistry and pressure dependent rates. A flux-based algorithm is used to decide which isomers to add to the network. Isomers added to the network are reacted to form other channels. When enough of the flux is accounted for by the isomers and bimolecular product channels, the generation process terminates to yield a comprehensive network. Network sensitivity analysis is applied to the network to identify the most important wells and barriers that can be refined using quantum chemistry calculations. Iterative refinement can then be used to achieve accuracy. The comprehensive network can be easily reduced using energy and flux-based algorithms to the most important channels and isomers in the network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Daylight simulation workflows incorporating measured bidirectional scattering distribution functions

Daylight predictions of architectural spaces depend on good estimates of light transfer through skylights, windows and other fenestration systems. For clear glazing and painted surfaces, parametric transmission and reflection models have proven adequate, but there are many cases where light-scattering, semi-specular shading and daylighting materials defy simple characterization. Something as commonplace as fabric roller shades and venetian blinds may turn daylight prediction into guesswork, and numerous advanced systems on the market tuned specifically to enhance daylight are not sufficiently characterized to distinguish their performance. In this paper, we describe new tools available to handle novel and specialized fabrics, materials, and devices using data-driven modelling of bi-directional scattering distribution functions (BSDFs). These representations are usually tabulated at constant or adjustable angular resolution for efficient point-in-time and annual daylight simulations. We describe a variety of BSDF simulation workflows, including some of the tools and methods that make advanced analysis possible, and highlight some of the current challenges. We conclude with a discussion of future work and how such data might be created and shared worldwide.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Soil porous microstructure control over soil organic matter mobility: A multimethod workflow for understanding chemistry-dependent organic matter binding in soil

Soil organic matter (SOM) has attracted a great deal of interest; particularly for its potential to mitigate human derived CO 2 emissions. Studies have demonstrated that SOM plays a critical role in carbon storage and CO 2 sequestration. However, the sorption properties of SOM, which influence its transport in pore water and stabilization within the soil, remain poorly understood. This study develops a workflow to: (1) examine compound-specific advective and diffusive transport and desorption behaviors, (2) quantify desorption rates through stop-flow and continuous-flow column experiments, and (3) evaluate the impact of soil microporosity on SOM mobility using high-resolution imaging and extractions. Intact core column experiments were conducted on Uncultivated (Natural) and Cultivated soil samples, both were arid soils, collected in Washington State. X-ray computed tomography was employed to measure porosity and pore connectivity, while Fourier-transform ion cyclotron resonance mass spectrometry was used to analyze SOM composition. The findings revealed that cultivation increased total carbon and nitrogen levels due to irrigation and fertilization, enhancing carbon capture potential in arid soils. In contrast, the Natural soil, characterized by higher porosity and connectivity, contained more oxidized carbon. Pore network analysis indicated that soil compaction in the Cultivated soil may lead to longer diffusion pathways, significantly influencing SOM transport and stability.

Hydraulic Properties↗

A deep redox proteome profiling workflow and its application to skeletal muscle of a Duchenne Muscular Dystrophy model

Perturbation to the redox state accompanies many diseases and its effects are viewed through oxidation of biomolecules, including proteins, lipids, and nucleic acids. The thiol groups of protein cysteine residues undergo an array of redox post-translational modifications (PTMs) that are important for regulation of protein and pathway function. To better understand what proteins are redox regulated following a perturbation, it is important to be able to comprehensively profile protein thiol oxidation at the proteome level. Herein, we report a deep redox proteome profiling workflow and demonstrate its application in measuring the changes in thiol oxidation along with global protein expression in skeletal muscle from mdx mice, a model of Duchenne Muscular Dystrophy (DMD). In depth coverage of the thiol proteome was achieved with >18,000 Cys sites from 5608 proteins in muscle being quantified. Compared to the control group, mdx mice exhibit markedly increased thiol oxidation, where ~2% shift in the median oxidation occupancy was observed. Further, pathway analysis for the redox data revealed that coagulation system and immune-related pathways were among the most susceptible to increased thiol oxidation in mdx mice, whereas protein abundance changes were more enriched in pathways associated with bioenergetics. This study illustrates the importance of deep redox profiling in gaining greater insight into oxidative stress regulation and pathways/processes that are perturbed in an oxidizing environment.

60 APPLIED LIFE SCIENCES↗

Experimental workflow to estimate model parameters for evaluating long term viscoelastic response of CO2 storage caprocks

Understanding the time-dependent behavior of reservoir and sealing formations is critical to assessing risks associated with geological carbon storage since time-dependent deformation strongly influences mechanical responses of some rock types. Many studies have evaluated the risk of CO2 leakage and induced seismicity by assuming poroelastic rheology in sealing formations. Few have considered viscoelastic or other time-dependent responses, where the existing literature adopts 1D models to represent long-term time-dependent responses. This is primarily because to date, the general form of a reasonable 3D time-dependent model for rocks remains unclear. In this paper, we address this unclear issue by proposing a new workflow to select constitutive modeling parameters to evaluate if a 3D viscoelastic model is reasonable using several-hour-long experimental data and a power-law response to extrapolate to the decades-long time frames of interest in geologic carbon storage. To provide experimental data, we conducted multi-level loading/unloading triaxial relaxation tests with four rock types. The experimental results showed that the maximum load relaxation observed is approximately 49%, with some rock types showing as little as 1.4%. Using a simple linear viscoelastic model, parameters were chosen such that a maximum deviation of 1.5 MPa in axial stress and 7 MPa in radial stress was attained with the extrapolated 30-year data. We found that a reasonable parameter range for the normalized elastic modulus is 0.1~2 for rocks with significant time-dependent responses and 0.01~0.06 for those with small time-dependent responses. No matter how significant time-dependent responses are for rocks considered, our results showed that the relaxation time has a general range of 1~10^10 s, whose time scale can be one or two orders higher than a time frame typically envisioned for CO2 injection projects.

Stress relaxation, 3D time-dependent model, viscoe↗

A generalized machine learning workflow to visualize mechanical discontinuity

Accurate detection and mapping of mechanical discontinuity in materials has widespread industrial and research applications. Herein, we developed a generalized machine-learning framework for visualizing single mechanical discontinuity embedded in material of any composition, velocity, density, porosity, and size with limited data. The proposed visualization of discontinuity requires accurate estimations of the length, location, and orientation of the embedded discontinuity by processing multipoint wave-transmission measurements. k-Wave simulator is used to create a large dataset of elastic waveforms recorded during multi-point wave-transmission measurements through materials containing single mechanical discontinuity. k-Wave simulator considers the wave attenuation, dispersion, and mode conversion in wave motion. Discrete wavelet transform (DWT) and statistical feature extraction are essential for data preprocessing prior to the data-driven model development. DWT also minimizes the effect of noise. Using hyper-parameter tuning and cross validation, gradient boosting regression can visualize the mechanical discontinuity with an accuracy of 0.85, in terms of coefficient of determination. A double-layered neural network-based regression has better performance with an accuracy of 0.95. Use of convolutional neural network converts the predictive task from a waveform processing to an image processing problem. Convolutional neural network achieved a generalization performance of 0.91. The proposed generalized workflow requires robust simulation of wave propagation, signal processing, feature engineering, and model evaluation. Sensors closest to the source and those located opposite the source are the most significant for the desired visualization. Notably, the sensors closest to the source capture the non-linear associations, whereas the sensor on the border opposite to the source capture the linear associations between the measured waveforms and the properties of the mechanical discontinuity.

42 ENGINEERING↗

Pynta-An Automated Workflow for Calculation of Surface and Gas-Surface Kinetics

Many important industrial processes rely on heterogeneous catalytic systems. However, given all possible catalysts and conditions of interest, it is impractical to optimize most systems experimentally. Automatically generated microkinetic models can be used to efficiently consider many catalysts and conditions. However, these microkinetic models require accurate estimation of many thermochemical and kinetic parameters. Manually calculating these parameters is tedious and error prone, involving many interconnected computations. Here, we present Pynta, a workflow software for automating the calculation of surface and gas–surface reactions. Pynta takes the reactants, products, and atom maps for the reactions of interest, generates sets of initial guesses for all species and saddle points, runs all optimizations, frequency, and IRC calculations, and computes the associated thermochemistry and rate coefficients. It is able to consider all unique adsorption configurations for both adsorbates and saddle points, allowing it to handle high index surfaces and bidentate species. Pynta implements a new saddle point guess generation method called harmonically forced saddle point searching (HFSP). HFSP defines harmonic potentials based on the optimized adsorbate geometries and which bonds are breaking and forming that allow initial placements to be optimized using the GFN1-xTB semiempirical method to create reliable saddle point guesses. This method is reaction class agnostic and fast, allowing Pynta to consider all possible adsorbate site placements efficiently. We demonstrate Pynta on 11 diverse reactions involving monodenate, bidentate, and gas-phase species, many distinct reaction classes, and both a low and a high index facet of Cu. Our results suggest that it is very important to consider reactions between adsorbates adsorbed in all unique configurations for interadsorbate group transfers and reactions on high index surfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Preprocessing Tool for Enhanced Ion Mobility–Mass Spectrometry-Based Omics Workflows

The ability to improve the data quality of ion mobility–mass spectrometry (IM-MS) measurements is of great importance for enabling modular and efficient computational workflows and gaining better qualitative and quantitative insights from complex biological and environmental samples. We developed the PNNL PreProcessor, a standalone and user-friendly software housing various algorithmic implementations to generate new MS-files with enhanced signal quality and in the same instrument format. Different experimental approaches are supported for IM-MS based on Drift-Tube (DT) and Structures for Lossless Ion Manipulations (SLIM), including liquid chromatography (LC) and infusion analyses. The algorithms extend the dynamic range of the detection system, while reducing file sizes for faster and memory-efficient downstream processing. Specifically, multidimensional smoothing improves peak shapes of poorly defined low-abundance signals, and saturation repair reconstructs the intensity profile of high-abundance peaks from various analyte types. Further, other functionalities are data compression and interpolation, IM demultiplexing, noise filtering by low intensity threshold and spike removal, and exporting of acquisition metadata. Several advantages of the tool are illustrated, including an increase of 19.4% in lipid annotations and a two-times faster processing of LC-DT IM-MS data-independent acquisition spectra from a complex lipid extract of a standard human plasma sample. The software is freely available at https://omics.pnl.gov/software/pnnl-preprocessor.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating Ultra-Coarse-Grained Protein Models into Accessible Workflows for Multiscale Molecular Dynamics

To capture protein conformational transitions using molecular dynamics (MD), several simulation resolutions covering different spatial and temporal scales are typically needed. All-atom (AA) simulations provide fine resolution, but are computationally infeasible for large systems over longer durations. Coarse-grained (CG) and ultra-coarse-grained (UCG) models have a lower resolution and computational cost while still being able to conserve essential protein features. Prior work on a Multiscale Machinelearned Modeling Infrastructure (MuMMI) combined both AA and CG simulations to study RAS-RAF protein interactions, leveraging CG models for longer time scales and using AA to investigate unusual conformations in greater detail. However, MuMMI is still resource-intensive, and this study aims to maximize exploration of the protein conformational space while reducing computational cost. In this paper, we build on prior work that integrates UCG models based on heterogeneous elastic network modeling (hENM) into the MuMMI workflow. We demonstrate that UCG models enable accurate sampling of protein conformations, focusing on simulating RAS-RAF protein interactions. Using higher-resolution CG Martini simulation data, we can automatically refine intramolecular interactions in UCG models. We present a scalable Python package that uses fluctuations observed in higher-resolution CG Martini simulations to estimate bond coefficients of the UCG model. We built novel machine learning-based backmapping methods to recover more detailed CG Martini structures from UCG structures, using diffusion models to learn the mapping between scales. Finally, we present UCG-mini-MuMMI, an accessible and less compute-intensive version of MuMMI as a resource for the scientific community. Incorporating UCG models into MD studies is applicable to a broad range of systems and proteins, and our study offers insights into the advantages and limitations of these methods.

Chemical structure↗

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY↗

Assessing the stability of Pd-exchanged sites in zeolites with the aid of a high throughput quantum chemistry workflow

Abstract Cation exchanged-zeolites are functional materials with a wide range of applications from catalysis to sorbents. They present a challenge for computational studies using density functional theory due to the numerous possible active sites. From Al configuration, to placement of extra framework cation(s), to potentially different oxidation states of the cation, accounting for all these possibilities is not trivial. To make the number of calculations more tractable, most studies focus on a few active sites. We attempt to go beyond these limitations by implementing a workflow for a high throughput screening, designed to systematize the problem and exhaustively search for feasible active sites. We use Pd-exchanged CHA and BEA to illustrate the approach. After conducting thousands of explicit DFT calculations, we identify the sites most favorable for the Pd cation and discuss the results in detail. The high throughput screening identifies many energetically favorable sites that are non-trivial. Lastly, we employ these results to examine NO adsorption in Pd-exchanged CHA, which is a promising passive NO x adsorbent (PNA) during the cold start of automobiles. The results shed light on critical active sites for NO x capture that were not previously studied.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

pyGWBSE: a high throughput workflow package for GW-BSE calculations

Abstract We develop an open-source python workflow package, py GWBSE to perform automated first-principles calculations within the GW-BSE (Bethe-Salpeter) framework. GW-BSE is a many body perturbation theory based approach to explore the quasiparticle (QP) and excitonic properties of materials. GW approximation accurately predicts bandgaps of materials by overcoming the bandgap underestimation issue of the more widely used density functional theory (DFT). BSE formalism produces absorption spectra directly comparable with experimental observations. py GWBSE package achieves complete automation of the entire multi-step GW-BSE computation, including the convergence tests of several parameters that are crucial for the accuracy of these calculations. py GWBSE is integrated with Wannier90 , to generate QP bandstructures, interpolated using the maximally-localized wannier functions. py GWBSE also enables the automated creation of databases of metadata and data, including QP and excitonic properties, which can be extremely useful for future material discovery studies in the field of ultra-wide bandgap semiconductors, electronics, photovoltaics, and photocatalysis.

36 MATERIALS SCIENCE↗

Developing a complete AI-accelerated workflow for superconductor discovery

The quest to identify new superconducting materials with enhanced properties is hindered by the prohibitive cost of computing electron-phonon spectral functions, severely limiting the materials space that can be explored. Here, we introduce a Bootstrapped Ensemble of Equivariant Graph Neural Networks (BEE-NET), a machine-learning model trained to predict the Eliashberg spectral function and superconducting critical temperature with a mean-absolute-error of 0.87 K relative to DFT-based Allen-Dynes calculations. Intriguingly, BEE-NET achieves a true-negative-rate of 99.4%, enabling highly efficient screening for the rare property of superconductivity. Integrated into a multi-stage, AI-accelerated discovery pipeline that incorporates elemental-substitution strategies and machine-learned interatomic potentials, our workflow reduced over 1.3 million candidate structures to 741 dynamically and thermodynamically stable compounds with DFT-confirmed T c > 5 K. We report the successful synthesis and experimental confirmation of superconductivity in two of these previously unreported compounds. This study establishes a data-driven framework that integrates machine learning, quantum calculations, and experiments to systematically accelerate superconductor discovery.

Gibson, Jason B. [Quantum Formatics, Cambridge, MA↗