Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

High-Throughput Automated Exploration of Phase Growth Behaviors in Quasi-2D Formamidinium Metal Halide Perovskites

Quasi-2D metal halide perovskites (MHPs) are an emerging material platform for sustainable functional optoelectronics, but the uncontrollable, broad phase distribution remains a critical challenge for applications. Nevertheless, the basic principles for controlling phases in quasi-2D MHPs remain poorly understood, due to the rapid crystallization kinetics during the conventional thin-film fabrication process. In this work, a high-throughput automated synthesis-characterization-analysis workflow is implemented to accelerate material exploration in formamidinium (FA)-based quasi-2D MHP compositional space, revealing the early-stage phase growth behaviors fundamentally determining the phase distributions. Upon comprehensive exploration with varying synthesis conditions including 2D:3D composition ratios, antisolvent injection rates, and temperatures in an automated synthesis-characterization platform, it is observed that the prominent n = 2 2D phase restricts the growth kinetics of 3D-like phases—α-FAPbI 3 MHPs with spacer-coordinated surface—across the MHP compositions. Thermal annealing is a critical step for proper phase growth, although it can lead to the emergence of unwanted local PbI 2 crystallites. Additionally, fundamental insights into the precursor chemistry associated with spacer-solvent interaction determining the quasi-2D MHP morphologies and microstructures are demonstrated. The high-throughput study provides comprehensive insights into the fundamental principles in quasi-2D MHP phase control, enabling new control of the functionalities in complex materials systems for sustainable device applications.

2D perovskites↗

How Climate and Data Quality Impact Photovoltaic Performance Loss Rate Estimations

Different data pipelines and statistical methods are applied to photovoltaic (PV) performance datasets to quantify the performance loss rate (PLR). Since the real values of PLR are unknown, a variety of unvalidated values are reported. As such, the PV industry commonly assumes PLR based on statistically extracted ranges from the literature. However, the accuracy and uncertainty of PLR depend on several parameters including seasonality, local climatic conditions, and the response of a particular PV technology. In addition, the specific data pipeline and statistical method used affect the accuracy and uncertainty. To provide insights, a framework of (≈200 million) synthetic simulations of PV performance datasets using data from different climates is developed. Time series with known PLR and data quality are synthesized, and large parametric studies are conducted to examine the accuracy and uncertainty of different statistical approaches over the contiguous US, with an emphasis on the publicly available and “standardized” library, RdTools . In the results, it is confirmed that PLRs from RdTools are unbiased on average, but the accuracy and uncertainty of individual PLR estimates vary with climate zone, data quality, PV technology, and choice of analysis workflow. Best practices and improvement recommendations based on the findings of this study are provided.

14 SOLAR ENERGY↗

Future Trends in Nuclear Physics Computing

In nuclear physics (NP) today the study of quarks, gluons and their strong interactions extends across a broad research program at a varied range of collaborative scales, from a few collaborators up to large experiments at scales comparable to those typical of high energy physics (HEP). Overall, the software and computing efforts vary accordingly, from pragmatic do-it-yourself approaches among a few, to substantial organized software and computing activities within large experiments. With new experiments starting up and on the horizon [1], and rapidly increasing data volumes [2, 3] and processing demands even at small experiments, the NP community has in recent years been thinking about the next generation of data processing and analysis workflows that will maximize the science output. One context for this discussion has been a series of workshops, “Future Trends in Nuclear Physics Computing” [4]. The most recent in this series took place in Fall 2020, organized by the authors together with colleagues. The workshop focused on identifying the unique aspects of software and computing in NP, and discussing how the NP community could strengthen common efforts and chart a path forward for the next decade, sure to be an exciting one with rich ongoing scientific programs at Brookhaven National Laboratory (BNL), Jefferson Lab (JLab), and other NP facilities, and culminating in datataking at the Electron-Ion Collider (EIC) [5,6,7] in the early 2030s. Without claiming to present a collective view from the workshop and discussions since—fortunately this is not expected of us in this opinion editorial—we offer here our reflections on the topic, informed by the workshop and the summary we authored with our colleagues [8], as well as discussions and developments in the eventful time since.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

nmRanalysis: An Open-Source Web Application for Semi-automated NMR Metabolite Profiling

Though data acquisition and initial signal pre-processing of nuclear magnetic resonance (NMR) spectra have achieved high degrees of automation, downstream processing - specifically the profiling of spectra - has bottlenecked the overall NMR analysis workflow. Several efforts have been made to mitigate this bottleneck, but these solutions often trade an increase in automation for limitations elsewhere. Here, in this technical note, we introduce nmRanalysis, a user-friendly web-application that integrates the strengths of existing profiling tools for a more automated profiling workflow. nmRa-nalysis additionally incorporates novel features, including a machine-learning-driven recommender system for me-tabolite identification, further increasing the utility of nmRanalysis over the individual tools that it incorporates.

Flores, Javier E. [Pacific Northwest National Labo↗

Predicting Trends in VOC Through Rapid, Multimodal Characterization of State-of-the-Art p-i-n Perovskite Devices

Perovskite photovoltaic technologies are approaching commercial deployment, yet single junction and tandem architectures both still have significant room to improve power conversion efficiency and stability. The ability to perform rapid screening of material quality after altering processing conditions is critical to accelerating the optimization and commercialization of perovskite-based technologies. Currently, researchers utilize a wide range of stand-alone metrology tools to isolate sources of power loss throughout a device stack, which can be slow and labor intensive. Here, we demonstrate the use of a multimodal metrology approach to rapidly determine the maximum achievable and predicted open circuit voltages of >100 perovskite devices during fabrication. Acquisition of these different data is facilitated by combining them into a single integrated measurement platform. We show that these data and automated analysis can be used to rapidly understand and ultimately predict quantitative trends in open circuit voltages of state-of-the-art device architectures. The data and automated analysis workflow presented provides a reliable approach to quickly identify absorber and charge transport layer combinations that can lead to improved open circuit voltages.

14 SOLAR ENERGY↗

evSeq: Cost-Effective Amplicon Sequencing of Every Variant in a Protein Library

Widespread availability of protein sequence-fitness data would revolutionize both our biochemical understanding of proteins and our ability to engineer them. Unfortunately, even though thousands of protein variants are generated and evaluated for fitness during a typical protein engineering campaign, most are never sequenced, leaving a wealth of potential sequence-fitness information untapped. Primarily, this is because sequencing is unnecessary for many protein engineering strategies; the added cost and effort of sequencing is thus unjustified. It also results from the fact that, even though many lower cost sequencing strategies have been developed, they often require at least some sequencing or computational resources, both of which can be barriers to access. In this work, we present every variant sequencing (evSeq), a method and collection of tools/standardized components for sequencing a variable region within every variant gene produced during a protein engineering campaign at a cost of cents per variant. evSeq was designed to democratize low-cost sequencing for protein engineers and, indeed, anyone interested in engineering biological systems. Execution of its wet-lab component is simple, requires no sequencing experience to perform, relies only on resources and services typically available to biology labs, and slots neatly into existing protein engineering workflows. Analysis of evSeq data is likewise made simple by its accompanying software (found at github.com/fhalab/evSeq, documentation at fhalab.github.io/evSeq), which can be run on a personal laptop and was designed to be accessible to users with no computational experience. Here, low-cost and easy to use, evSeq makes collection of extensive protein variant sequence-fitness data practical.

59 BASIC BIOLOGICAL SCIENCES↗

Guiding the choice of informatics software and tools for lipidomics research applications

Progress in mass spectrometry lipidomics has led to a rapid proliferation of studies across biology and biomedicine. These generate extremely large raw datasets requiring sophisticated solutions to support automated data processing. To address this, numerous software tools have been developed and tailored for specific tasks. However, for researchers, deciding which approach best suits their application relies on ad hoc testing, which is inefficient and time consuming. Here we first review the data processing pipeline, summarizing the scope of available tools. Next, to support researchers, LIPID MAPS provides an interactive online portal listing open-access tools with a graphical user interface. This guides users towards appropriate solutions within major areas in data processing, including (1) lipid-oriented databases, (2) mass spectrometry data repositories, (3) analysis of targeted lipidomics datasets, (4) lipid identification and (5) quantification from untargeted lipidomics datasets, (6) statistical analysis and visualization, and (7) data integration solutions. Detailed descriptions of functions and requirements are provided to guide customized data analysis workflows.

59 BASIC BIOLOGICAL SCIENCES↗

The fast camera (Fastcam) imaging diagnostic systems on the DIII-D tokamak

Two camera systems are installed on the DIII-D tokamak at the toroidal positions of 90° (90° system) and 225° (225° system), respectively. The cameras have two types of relay optics, namely, a coherent optical fiber bundle and a periscope system. The periscope system provides absolute intensity calibration stability while sacrificing resolution (10 lp/mm), while the fiber system provides high resolution (16 lp/mm) while sacrificing calibration stability. The periscope is available only for the 90° system. The optics of the 225° system were designed for view stability, repeatability, and easy maintenance. The cameras are located inside optimized neutron, x ray and magnetic shielding in order to reduce electronics damage, reboots, and magnetic and neutron interference, increasing the overall system reliability. An automated filter wheel, providing remote filter change, allows for remote wavelength selection. A software suite automates camera acquisition and data storage, allowing for remote operation and reduced operator involvement. System metadata is used to streamline the data analysis workflow, particularly for intensity calibration. Here, the spatial calibration uses multiple observable wall features, resulting in a reconstruction accuracy ≤2 cm.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

WormBase in 2022—data, processes, and tools for analyzing Caenorhabditis elegans

WormBase (www.wormbase.org) is the central repository for the genetics and genomics of the nematode Caenorhabditis elegans. We provide the research community with data and tools to facilitate the use of C. elegans and related nematodes as model organisms for studying human health, development, and many aspects of fundamental biology. Throughout our 22-year history, we have continued to evolve to reflect progress and innovation in the science and technologies involved in the study of C. elegans. We strive to incorporate new data types and richer data sets, and to provide integrated displays and services that avail the knowledge generated by the published nematode genetics literature. Here, we provide a broad overview of the current state of WormBase in terms of data type, curation workflows, analysis, and tools, including exciting new advances for analysis of single-cell data, text mining and visualization, and the new community collaboration forum. Concurrently, we continue the integration and harmonization of infrastructure, processes, and tools with the Alliance of Genome Resources, of which WormBase is a founding member.

59 BASIC BIOLOGICAL SCIENCES↗

MapsTorch : automatic differentiation for X-ray fluorescence data analysis

X-ray fluorescence (XRF) is a popular spectroscopy technique for elemental analysis. Spectrum fitting and parameter tuning are at the core of XRF analysis and are conventionally manually intensive, especially for synchrotron experiments involving large amounts of diverse samples. This work introduces the automatic differentiation (AD) technique to XRF and an open-source package called MapsTorch. By transforming an analytical model of the XRF spectrum into a differentiable computation graph with AD, MapsTorch enables robust optimization of parameters and elemental intensities. We evaluate MapsTorch by conducting computational experiments on a large number of historical synchrotron XRF datasets and compare its performance with the currently practiced fitting tool NLopt. The results show that MapsTorch consistently achieves high-quality fits and often leads to better fitting quality than NLopt, particularly in tasks such as initial spectrum fitting and elemental intensity refinement. The robust performance of MapsTorch paves the way for developing automated and high-throughput XRF data analysis workflows to handle the increasing data volumes expected from next-generation synchrotron facilities.

X-ray fluorescence↗

Performance of the reconstruction of large impact parameter tracks in the inner detector of ATLAS

Searches for long-lived particles (LLPs) are among the most promising avenues for discovering physics beyond the Standard Model at the Large Hadron Collider (LHC). However, displaced signatures are notoriously difficult to identify due to their ability to evade standard object reconstruction strategies. In particular, the ATLAS track reconstruction applies strict pointing requirements which limit sensitivity to charged particles originating far from the primary interaction point. To recover efficiency for LLPs decaying within the tracking detector volume, the ATLAS Collaboration employs a dedicated large-radius tracking (LRT) pass with loosened pointing requirements. During Run 2 of the LHC, the LRT implementation produced many incorrectly reconstructed tracks and was therefore only deployed in small subsets of events. In preparation for LHC Run 3, ATLAS has significantly improved both standard and large-radius track reconstruction performance, allowing for LRT to run in all events. This development greatly expands the potential phase-space of LLP searches and streamlines LLP analysis workflows. This paper will highlight the above achievement and report on the readiness of the ATLAS detector for track-based LLP searches in Run 3.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data

Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.

Wang, Daoce [University of Nebraska, Omaha]↗

Whole Energy Homes

This repository contains analysis workflows, associated python packages, and documentation of the Whole Energy Homes study. A reproducible Snakemake workflow for data processing, feature extraction, and modeling using the weh python package. Built with Copier from the able-workflow-copier template.

Pathak, Maharshi [Northeastern Univ., Boston, MA (↗

NOODLES [SWR-22-78]

NOODLES is a protocol specification for collaborative visualization. It allows software tools of any type to participate in an analysis or visualization session. Use cases include, but are not limited to, distributed analysis, workflow interoperability, computational steering, etc.

Brunhart-Lupo, Nicholas↗

ESS-DIVE Reporting Format for Amplicon Abundance Table

While standardized sequencing data is available in public repositories and efforts such as MIxS for common sample collection and processing metadata are well established, the lack of common bioinformatic processing metadata has hindered the ability to do large-scale metaanalyses and the potential for data re-use by non-experts such as ecosystem, watershed, or earth system modelers. To address this need for Department of Energy researchers, we have developed an amplicon reporting format which captures both sample preparation and bioinformatic processing metadata and stores processed amplicon data as a paired abundance table and sequencing file to maximize the potential for re-use of these data. To aid in the adoption of accessible and reproducible analysis workflows, this reporting format was developed in concert with amplicon functionality within the Department of Energy’s Systems Biology Knowledgebase (KBase) to ensure common data and metadata requirements and facilitate seamless transfer between these platforms.This dataset contains support documentation for the amplicon reporting format (README.md and instructions.md), templates for both bioinformatic and sequencing metadata (amplicon_bioinformatic_metadata_template_2021_10_03.csv and amplicon_sequencing_metadata_template_2021_10_03.csv), a crosswalk indicating how this reporting format relates to the current MIxS format (ESSDIVE-MIxS_crosswalk.csv), a list of available instrument terms (amplicon_seq_instrument_terms_2021_10_03.csv), a map between QIIME2 parameter settings and metadata fields (amplicon_qiime2_plugin_metadata_map.csv), a data dictionary (amplicon_CSV_dd.csv), and file-level metadata (amplicon_FLMD.csv).

54 ENVIRONMENTAL SCIENCES↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Augmented Human Analysis (AHA)

Radio frequency (RF) signal monitoring generally emphasizes intentionally generated signals, such as WiFi, Bluetooth, or cellular transmissions. However, electronic devices also produce unintended radiated emissions (UREs), which could also be useful in RF spectrum analysis. In either case, deriving intelligence from RF signals is typically a human-intensive process requiring significant domain knowledge. In the Augmented Human Analysis (AHA) project, we investigate the utility of dimensionally aligned signal projection (DASP) and machine learning (ML) algorithms for accelerating RF analysis workflows. We find that while DASP algorithms can indeed highlight signal characteristics relevant for classification tasks, the choice of algorithmic hyperparameters greatly affects performance. To address this challenge, we evaluate the quality of DASP outputs using the silhouette score, which measures how well data points cluster; high silhouette scores indicate good clustering, and thus good hyperparameter values. This approach is critical for machine learning pipelines as the DASP parameters cannot be directly optimized during model training. By identifying good DASP parameters, and thus good DASP outputs, as a preprocessing step, we can decrease the amount of effort required for downstream ML model training. We demonstrate our workflow using a dataset of UREs from common household devices, showing that even without the aid of ML, proper selection of DASP parameters enables clustering by device type.

42 ENGINEERING↗