Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data reduction pipelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

67 records · Page 4

Archetype-based Redshift Estimation for the Dark Energy Spectroscopic Instrument Survey

We present a computationally efficient galaxy archetype-based redshift estimation and spectral classification method for the Dark Energy Survey Instrument (DESI) survey. The DESI survey currently relies on a redshift fitter and spectral classifier using a linear combination of principal component analysis–derived templates, which is very efficient in processing large volumes of DESI spectra within a short time frame. However, this method occasionally yields unphysical model fits for galaxies and fails to adequately absorb calibration errors that may still be occasionally visible in the reduced spectra. Our proposed approach improves upon this existing method by refitting the spectra with carefully generated physical galaxy archetypes combined with additional terms designed to absorb data reduction defects and provide more physical models to the DESI spectra. We test our method on an extensive data set derived from the survey validation (SV) and Year 1 (Y1) data of DESI. Our findings indicate that the new method delivers marginally better redshift success for SV tiles while reducing catastrophic redshift failure by 10%–30%. At the same time, results from millions of targets from the main survey show that our model has relatively higher redshift success and purity rates (0.5%–0.8% higher) for galaxy targets while having similar success for QSOs. These improvements also demonstrate that the main DESI redshift pipeline is generally robust. Additionally, it reduces the false-positive redshift estimation by 5%–40% for sky fibers. We also discuss the generic nature of our method and how it can be extended to other large spectroscopic surveys, along with possible future improvements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Validation of the DESI DR2 measurements of baryon acoustic oscillations from galaxies and quasars

The Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) galaxy and quasar clustering data represents a significant expansion of data from Data Release 1 (DR1), providing improved statistical precision in baryon acoustic oscillation (BAO) constraints across multiple tracers, including bright galaxies, luminous red galaxies, emission line galaxies, and quasars. In this paper, we validate the BAO analysis of DR2. We present the results of robustness tests on the blinded DR2 data and, after unblinding, consistency checks on the unblinded DR2 data. All results are compared with those obtained from a suite of mock catalogs that replicate the selection and clustering properties of the DR2 sample. We confirm the consistency of DR2 BAO measurements with DR1 while achieving a reduction in statistical uncertainties due to the increased survey volume and completeness. The combined BAO precision, including both statistical and systematic errors, improves from ∼0.52% in DR1 to 0.30% in DR2—a factor of 1.7 gain. We assess the impact of analysis choices, including different data vectors (correlation function vs power spectrum), modeling approaches and systematics treatments, and an assumption of the Gaussian likelihood, finding that our BAO constraints are stable across these variations and assumptions with a few minor refinements to the baseline setup of the DR1 BAO analysis. We summarize a series of pre-unblinding tests that confirmed the readiness of our analysis pipeline, the final systematic errors, and the DR2 BAO analysis baseline. The successful completion of these tests led to the unblinding of the DR2 BAO measurements, ultimately leading to the DESI DR2 cosmological analysis, with their implications for the expansion history of the Universe and the nature of dark energy presented in the DESI key paper (companion paper).

79 ASTRONOMY AND ASTROPHYSICS↗

Leveraging unlabeled SEM datasets with self-supervised learning for enhanced particle segmentation

Scanning Electron Microscopes (SEMs) are widely used in experimental science laboratories, often requiring cumbersome and repetitive user analysis. Automating SEM image analysis processes is highly desirable to address this challenge. In particle sample analysis, Machine Learning (ML) has emerged as the most effective approach for particle segmentation. However, the time-intensive process of manually annotating thousands of SEM images limits the applicability of supervised learning approaches. Self-Supervised Learning (SSL) offers a promising alternative by enabling knowledge extraction from raw, unlabeled data. This study presents a framework for evaluating SSL techniques in SEM image analysis, focusing on novel methods leveraging the ConvNeXtV2 architecture for particle detection. A dataset comprising 25,000 SEM images is curated to benchmark these proposed SSL methods. The results demonstrate that ConvNeXtV2 models, with varying parameter counts, consistently outperform other techniques in particle detection across different length scales, achieving up to a 34% reduction in relative error compared to established SSL methods. Furthermore, an ablation study explores the relationship between dataset size and SSL performance, providing actionable insights for practitioners regarding model selection and resource efficiency. This research advances the integration of SSL into autonomous analysis pipelines and supports its application in accelerating materials science discovery.

Rettenberger, Luca↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

Data for Impact of Non-Irrigation on 1G and 2G Bioethanol Potential of Oilcane Feedstock: A Field to Fuel Pipeline Study

This study evaluates the bioethanol potential in response to irrigation (IR) and non-irrigation (NIR) of oilcane (OC) during a seasonal drought prior harvest. The juice was extracted through mechanical pressing of stems and fermented by Ethanol Red® yeast to produce first-generation bioethanol. Hydrothermal pretreatment followed by enzymatic hydrolysis of bagasse was performed to produce monomeric sugars from structural carbohydrates. The hydrolysates were fermented with engineered yeast for second-generation bioethanol production. The irrigated oilcane juice (276.3 ± 8.9 g/L) constitutes higher sugar concentrations than non-irrigated oilcane juice (236.5 ± 2.2 g/L). The enzymatic hydrolysis of IR-OC and NIR-OC pretreated bagasse yielded similar concentrations of 247.5 ± 2.22 and 249.7 ± 4.98 g/L fermentable sugars. Industry-relevant bioethanol titers of ≥99 g/L and ≥75 g/L were achieved from juice and hydrolysates, respectively. Therefore, the non-irrigation regime did not impact the 1G and 2G bioethanol titers. However, the overall bioethanol yield can be lower due to the reduction of stem yield (8 %) per hectare.

Biomass Analytics↗

Uncovering Structure–Conductivity Relationships in Anion Exchange Membranes (AEMs) Using Interpretable Machine Learning

Anion exchange membranes (AEMs) play a vital role in the performance of water electrolyzers and fuel cells, yet their discovery and optimization remain challenging due to the complexity of structure–property relationships. In this study, we introduce a machine learning framework that leverages conditional graph neural networks (cGNNs) and descriptor-based models and a hybrid graph neural network (HGARE) to predict and interpret ionic conductivity. The descriptor-based pipeline employs principal component analysis (PCA), ablation, and SHAP analysis to identify factors governing anion conductivity, revealing electronic, topological, and compositional descriptors as key contributors. Beyond prediction, dimensionality reduction and clustering are performed by employing t-SNE and KMeans as well as SOM, which reveal distinct membranes clusters, some of which were enriched with high anion conductivity. Among graph-based approaches, the graph convolutional (GCN) achieved strong predictive performance, while the Hybrid Graph Autoencoder-Regressor Ensemble (HGARE) achieved the highest accuracy. Additionally, atom-level saliency maps from GCN provide spatial explanations for conductive behavior, revealing the importance of polarizable and flexible regions. This work contributes to the accelerated and data-driven design of high-performance AEMs.

Naghshnejad, Pegah [Department of Chemical Enginee↗

Enabling and Enhancing Space Mission Success and Reduction of Risk through the Application of an Integrated Data Architecture

The engineering phases of design, development, test, and evaluation (DDT and E) and subsequent planning, preparation, and operation (Ops) of space vehicles in a complex and distributed environment requires massive and continuous flows of information across the enterprise and across temporal stages of the vehicle lifecycle. The resulting capabilities at each subsequent stage depend in part on the capture, preparation, storage, and subsequent provision of information from prior stages. The United States National Aeronautics and Space Administration (NASA) is currently designing a fleet of new vehicles that will replace the Space Shuttle and expand space operations and exploration capabilities. This includes the 2 stage human rated lift vehicle Ares 1 and its associated crew vehicle the Orion, and a service module; the heavy lift cargo vehicle, Ares 5, and an associated cargo stage known as the Earth Departure Stage; and a Lunar Lander vehicle that contains a descent stage, and ascent stage, and a habitation module. A variety of concurrent assorted ground operations infrastructure including software and facilities are also being developed, assorted technology and assembly designs and development for equipment such as EVA suits, life support systems, command and control technologies are also in the pipeline. The development is occurring in a distributed manner, with project deliverables being contributed by a large and diverse assortment of vendors and most space faring nations. Critical information about all of the components, software, and procedures must be shared during the DDT and E phases and then made readily available to the mission operations staff for access during the planning, preparation, and operations phases, and also need to be readily available for system to system interactions. The Constellation Data Systems Project (CxDS) is identifying the needs, and designing and deploying systems and processes to support these needs. This paper details the steps and processes that NASA is applying within the Constellation Program to manage this data and information, and to insure that the correct information is available, correctly annotated, and can be provisioned digitally to enhance response times, and support engineering analysis and anomaly resolution.

Brummett, Robert C.↗

Cryptic cycling by electroactive bacterioplankton in Trout Bog Lake

The potential for extracellular electron transfer (EET) is a prevailing genomic feature of humic lake bacterioplankton. However, there has been little evidence for the substantial ecological contribution predicted by genetics. We hypothesized that anoxygenic phototrophic electrotrophs and accompanying heterotrophic electrogens cycle dissolved organic matter (DOM) between oxidized and reduced states. We predicted that such bacterioplankton would exhibit diel-scale oscillations due to the light dependency of photosynthesis. Using Trout Bog Lake in Wisconsin, USA, as our model ecosystem, we profiled the water column with depth-discrete metagenomic, physiochemical, and electrochemical analyses. We observed variation in oxidation reduction potential (ORP) in response to sunlight, initiating at depths populated by anoxygenic phototrophs with EET genes. We developed an automated buoy to measure electric current flow between many pairs of electrodes simultaneously, observing correlation in electron consumption to sunlight. Our results, combined with published metatranscriptomic analysis, indicate the occurrence of electron cycling between phototrophic oxidation (electrotrophic metabolism) by Chlorobium and anaerobic respiration (electrogenic metabolism) by Geothrix, involving DOM. We also repeatedly observed gradual seasonal increases in hypolimnion ORP throughout summer. These diel and seasonal patterns imply that electroactive DOM mediates the ecology of electroactive bacteria in lakes, controlling humic lake methane emissions.IMPORTANCEWe investigated the physical, chemical, and redox characteristics of a bog lake and electrodes hung therein to test the hypothesis that dissolved organic matter is being cycled between oxidized and reduced states by electroactive bacterioplankton powered by phototrophy. To do so, we performed field-based analyses on multiple timescales using both established and novel instrumentation. We paired these analyses with recently developed bioinformatics pipelines for metagenomics data to investigate genes that enable electroactive metabolism and accompanying metabolisms. Our results are consistent with our hypothesis and yet upend some of our other expectations. Our findings have implications for understanding greenhouse gas emissions from lakes, including electroactivity as an integral part of lake metabolism throughout more of the anoxic parts of lakes and for a longer portion of the summer than expected. Our results also give a sense of what electroactivity occurs at given depths and provide a strong basis for future studies.

carbon emissions↗

Transplatformer: translating toxicogenomic profiles between generations of platforms

Background Transcriptomic profiling technologies have advanced the analysis of biological and toxicological responses. However, substantial differences in probe design, dynamic range, gene coverage, and preprocessing pipelines across platforms introduce artifacts that limit cross-study integration and hinder the reuse of historical datasets. We aim to develop computational methods for accurate cross-platform translation to maximize the value of legacy resources. Results We present TransPlatformer a deep learning framework for translating gene expression profiles across heterogeneous toxicogenomics platforms. TransPlatformer employs a novel attention-based architecture to map high-dimensional fold-change vectors from legacy microarray technologies to current platforms. Models are trained and evaluated using DrugMatrix, spanning three technological generations. We investigate mixed-tissue, single-tissue, and cross-tissue training paradigms and benchmark performance against multilayer perceptron and matrix-completion baselines. In mixed-tissue training, TransPlatformer achieves a greater than 50% reduction in mean absolute error (0.043 vs. 0.09) and nearly doubles Pearson correlation ( ≈ 0.71 vs. 0.37) relative to baseline methods. Importantly, TransPlatformer preserves rare but biologically meaningful over- and under-expressed signals, with mean absolute error below 0.22. Single-tissue models yield further improvements for well-represented organs, such as a 10% reduction in liver mean absolute error, while underscoring the need for data augmentation strategies in low-sample tissues.ra Conclusions TransPlatformer provides an effective and scalable computational solution for cross-platform transcriptomic translation. By enabling biologically faithful harmonization of gene expression data, the proposed approach facilitates the reuse of legacy toxicogenomics datasets, enhances downstream biomarker discovery, and supports more reproducible predictive modeling in toxicology.

59 BASIC BIOLOGICAL SCIENCES↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗

Growth Curve Parameterization of Metabolic Activity of Yeast Cells for BioSentinel

The goal of the BioSentinel small satellite payload is to measure the effect of deep space radiation on the growth and metabolic activity of yeast cells. Raw test data is generated by fluidics cards containing yeast cells rehydrated at different periods, with metabolic activity measured by the reduction of alamarBlue. Each card well has a sensor array that measures the amount of red, green, and infrared light transmitted through the yeast culture. This illumination data is then converted to absorbance values, which are further converted into concentrations. The ultimate objective is to convert these concentrations into biologically-relevant metrics that can be compared against one another to determine changes due to differential radiation exposure. Beginning with IR absorbance data (corresponding to cell density) from ground studies, three parameters from a sigmoidal growth curve were extracted and analyzed: 𝜆 (lag phase), 𝜇 (max growth rate), and A (max cell growth). The data was fit to the Gompertz model of microbial growth using non-linear regression (Minitab), as the fit error was reduced compared to the simpler logistic growth curve. Graphs showed that the data contained a discrepancy (drift) in the lag phase that is attributable to a slow, constant loss of moisture. Correcting this discrepancy by fitting the first 25 hours of the data to a power function and subtracting these values from the absorbance readings obtained a better statistical fit to the growth curve in the lag phase. A power fit was selected over a linear fit because it reflected the effects of constant volume loss. This correction to the BioSentinel data analysis pipeline will enable quantitative statistical analysis of the effect of different levels of deep space radiation on yeast cells. Future work includes automation of drift correction and curve modeling to extract these parameters directly from data.

Growth Curve↗

NASA Tech Briefs, December 2005

Topics covered include: Video Mosaicking for Inspection of Gas Pipelines; Shuttle-Data-Tape XML Translator; Highly Reliable, High-Speed, Unidirectional Serial Data Links; Data-Analysis System for Entry, Descent, and Landing; Hybrid UV Imager Containing Face-Up AlGaN/GaN Photodiodes; Multiple Embedded Processors for Fault-Tolerant Computing; Hybrid Power Management; Magnetometer Based on Optoelectronic Microwave Oscillator; Program Predicts Time Courses of Human/ Computer Interactions; Chimera Grid Tools; Astronomer's Proposal Tool; Conservative Patch Algorithm and Mesh Sequencing for PAB3D; Fitting Nonlinear Curves by Use of Optimization Techniques; Tool for Viewing Faults Under Terrain; Automated Synthesis of Long Communication Delays for Testing; Solving Nonlinear Euler Equations With Arbitrary Accuracy; Self-Organizing-Map Program for Analyzing Multivariate Data; Tool for Sizing Analysis of the Advanced Life Support System; Control Software for a High-Performance Telerobot; Java Radar Analysis Tool; Architecture for Verifiable Software; Tool for Ranking Research Options; Enhanced, Partially Redundant Emergency Notification System; Close-Call Action Log Form; Task Description Language; Improved Small-Particle Powders for Plasma Spraying; Bonding-Compatible Corrosion Inhibitor for Rinsing Metals; Wipes, Coatings, and Patches for Detecting Hydrazines; Rotating Vessels for Growing Protein Crystals; Oscillating-Linear-Drive Vacuum Compressor for CO2; Mechanically Biased, Hinged Pairs of Piezoelectric Benders; Apparatus for Precise Indium-Bump Bonding of Microchips; Radiation Dosimetry via Automated Fluorescence Microscopy; Multistage Magnetic Separator of Cells and Proteins; Elastic-Tether Suits for Artificial Gravity and Exercise; Multichannel Brain-Signal-Amplifying and Digitizing System; Ester-Based Electrolytes for Low-Temperature Li-Ion Cells; Hygrometer for Detecting Water in Partially Enclosed Volumes; Radio-Frequency Plasma Cleaning of a Penning Malmberg Trap; Reduction of Flap Side Edge Noise - the Blowing Flap; and Preventing Accidental Ignition of Upper-Stage Rocket Motors.

Source record↗

Navigating the Noise: Bringing Clarity to ML Parameterization Design With O $\boldsymbol{\mathcal{O}}$(100) Ensembles

Abstract Machine‐learning (ML) parameterizations of subgrid processes (here of turbulence, convection, and radiation) may one day replace conventional parameterizations by emulating high‐resolution physics without the cost of explicit simulation. However, uncertainty about the relationship between offline and online performance (i.e., when integrated with a large‐scale general circulation model) hinders their development. Much of this uncertainty stems from limited sampling of the noisy, emergent effects of upstream ML design decisions on downstream online hybrid simulation. Our work rectifies the sampling issue via the construction of a semi‐automated, end‐to‐end pipeline for size ensembles of hybrid simulations, revealing important nuances in how systematic reductions in offline error manifest in changes to online error and online stability. For example, removing dropout and switching from a Mean Squared Error to a Mean Absolute Error loss both reduce offline error, but they have opposite effects on online error and online stability. Other design decisions, like incorporating memory, converting moisture input from specific humidity to relative humidity, using batch normalization, and training on multiple climates do not come with any such compromises. Finally, we show that ensemble sizes of may be necessary to reliably detect causally relevant differences online. By enabling rapid online experimentation at scale, we can empirically settle debates regarding subgrid ML parameterization design that would have otherwise remained unresolved in the noise.

Lin, Jerry [Department of Earth System Sciences Un↗