Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data transfer pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

VLSI high speed packet processor

The Goddard Space Flight Center Mission Operations and Data Systems Directorate has developed a packet processor card utilizing semicustom very large scale integration (VLSI) devices, microprocessors, and programmable gate arrays to support the implementation of multichannel telemetry data capture systems. This card will receive synchronized error corrected telemetry transfer frames and output annotated application packets derived from this data. An adaptable format capability is provided by the programmability of three microprocessors while the throughput capability of the packet processor is achieved by a data pipeline consisting of two separate RAM systems controlled by specially designed semicustom VLSI logic.

Grebowsky, Gerald J.↗

Workflow for High-throughput Screening of Enzyme Mutant Libraries Using Matrix-assisted Laser Desorption/Ionization Mass Spectrometry Analysis of Escherichia coli Colonies

High-throughput molecular screening of microbial colonies and DNA libraries are critical procedures that enable applications such as directed evolution, functional genomics, microbial identification, and creation of engineered microbial strains to produce high-value molecules. A promising chemical screening approach is the measurement of products directly from microbial colonies via optically guided matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). Measuring the compounds from microbial colonies bypasses liquid culture with a screen that takes approximately 5 s per sample. We describe a protocol combining a dedicated informatics pipeline and sample preparation method that can prepare up to 3,000 colonies in under 3 h. The screening protocol starts from colonies grown on Petri dishes and then transferred onto MALDI plates via imprinting. The target plate with the colonies is imaged by a flatbed scanner and the colonies are located via custom software. The target plate is coated with MALDI matrix, MALDI-MS analyzes the colony locations, and data analysis enables the determination of colonies with the desired biochemical properties. This workflow screens thousands of colonies per day without requiring additional automation. The wide chemical coverage and the high sensitivity of MALDI-MS enable diverse screening projects such as modifying enzymes and functional genomics surveys of gene activation/inhibition libraries.

Choe, Kisurb↗

Northeast Regional Planetary Data Center

In 1980, the Northeast Planetary Data Center (NEPDC) was established with Tim Mutch as its Director. The Center was originally located in the Sciences Library due to space limitations but moved to the Lincoln Field Building in 1983 where it could serve the Planetary Group and outside visitors more effectively. In 1984 Dr. Peter Schultz moved to Brown University and became its Director after serving in a similar capacity at the Lunar and Planetary Institute since 1976. Debbie Glavin has served as the Data Center Coordinator since 1982. Initially the NEPDC was build around Tim Mutch's research collection of Lunar Orbiter and Mariner 9 images with only partial sets of Apollo and Viking materials. Its collection was broadened and deepened as the Director (PHS) searched for materials to fill in gaps. Two important acquisitions included the transfer of a Viking collection from a previous PI in Tucson and the donation of surplused lunar materials (Apollo) from the USGS/Menlo Park prior to its building being torn down. Later additions included the pipeline of distributed materials such as the Viking photomosaic series and certain Magellan products. Not all materials sent to Brown, however, found their way to the Data Center, e.g., Voyager prints and negatives. In addition to the NEPDC, the planetary research collection is separately maintained in conjunction with past and ongoing mission activities. These materials (e.g., Viking, Magellan, Galileo, MGS mission products) are housed elsewhere and maintained independently from the NEPDC. They are unavailable to other researchers, educators, and general public. Consequently, the NEPDC represents the only generally accessible reference collection for use by researchers, students, faculty, educators, and general public in the Northeast corridor.

Schultz, Peter H.↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDRs success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

accessibility↗

Effect of Drying on Corrosion Mitigation of Hanford Transfer Lines

Radioactive waste is stored in underground, carbon-steel double-shell tanks at the Department of Energy Hanford site. The waste is transferred between the tanks and other assets using the transfer lines spanning throughout the various tank farms at Hanford. The transfer lines consist of a pipe-in-pipe design, small diameter pipes, and are not piggable. Recent inspection data of the transfer lines have shown areas with corrosion on both interior of the encasements and exterior of the primary pipes, with nearly 50 percent wall loss on the primary pipes and nearly 25% wall loss on the encasement pipes due to pitting corrosion. The visual inspections of the transfer lines have shown presence of corrosion products near the pipeline risers and beyond. It has been hypothesized that the corrosion is predominantly due to the high humidity conditions and in some cases is driven by the presence of residual hydrotest water in the encasement and the associated contact with the safety significant primary pipe. Therefore, drying of the transfer lines could lead to corrosion mitigation. Experimental studies are being conducted to understand the effect of environmental conditions, especially, relative humidity and temperature, on transfer line grade carbon steel corrosion and on mitigating corrosion. The experimental conditions are selected based on the seasonal temperature changes, and relative humidity conditions ranging from 30 to 100 percent. The experimental data will be used as guidance for maintaining a dry environment that will help mitigate the transfer-line corrosion caused by the high humidity conditions.

Shukla, Pavan K.↗

Evaluation of High Level Waste Sludge Processing Behavior

The U.S. Department of Energy’s (DOE) Hanford Site has 177 underground storage tanks that contain wastes from past nuclear fuel reprocessing and waste-management operations. Over 20% of this waste is in the form of an insoluble sludge that will require slurry modification before its transfer to the Waste Treatment and Immobilization Plant (WTP). Specific WTP acceptance criteria for waste feed delivery describe the physical and chemical characteristics of the waste that must be met before the waste is transferred to the WTP. One challenging requirement relates to the undissolved solids (UDS) composition in a waste feed because the waste contains solid particles that settle, and their concentration and relative proportion can change during the transfer of the waste in individual batches. A key uncertainty is the ability to transfer and mix wastes with large variations in UDS concentrations and resulting settling rates. To address this uncertainty, a number of small scale mixing and settling tests have been conducted to determine the mobilization performance of variable chemistry simulants. Comparison of the size and density of the particulate for each simulant to that of southeast area Hanford sludge was made using metrics for particle mobilization, suspension, settling, and pipeline transfer where dependance on particle size and density may be different, including: 1. Settling velocity, 2. Critical shear stress for erosion, 3. Just-suspended impeller speed, and 4. Pipeline critical transport velocity. Existing high-level waste sludge data has shown the effect that increasing Al concentration has on resulting settled solids. This differential settling of particles in the sludge has the possibility of resulting in solids segregation during feed preparation and uneven particle distribution during pipeline transportation or mixer jet pump operations. Understanding the predictive capabilities of HLW solids settling and transport as well as potential remedies for addressing disparate sludge behaviors can help provide technical guidance during HLW flowsheet planning.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

DEVELOP’s Approach to Experiential Learning

The NASA DEVELOP Program addresses environmental decision making needs and geoscience workforce development through 10-week feasibility studies that apply Earth observations to environmental issues at hand. The program builds capacity to use geospatial information in both its participants (students, recent graduates, early career professionals, and transitioning career professionals) and partner organizations (federal agencies, state & local governments, non-profits, and private industry). This is accomplished through a structured project execution model that provides opportunities for participants to have autonomy, learn “on the job,” and gain new skillsets for working with remote sensing data. A pipeline of leadership positions enhances opportunities for individuals engaged in the program to get hands-on experience conducting data analyses, communicating their work, leading technical projects, and building their knowledge bank of Earth-observing satellite capabilities. These skillsets and knowledge are then transferred to partners through the projects. This panel contribution will introduce the DEVELOP model, highlight experiences of participants, and share the program’s insights into good practices for effective experiential learning.

Capacity Building↗

Improving the Accessibility and Usability of Geothermal Information with Data Lakes and Data Pipelines on the Geothermal Data Repository: Preprint

The Geothermal Data Repository (GDR) provides universal access to data and information resulting from research and development activities funded by the Department of Energy (DOE). The GDR has extended this universal access to big data through integration with data lakes developed by the Open Energy Data Initiative (OEDI). Previously, large datasets such as seismic waveform or distributed acoustic sensing (DAS) data could only be accessed by institutions with high performance data storage and compute capabilities, effectively limiting the accessibility of big data to national labs, larger universities, and major corporations. Moreover, the time and resources needed to transport big data and configure them can produce additional barriers to use. Many of the standard formats used for structured data models (also known as content models) are incapable of handling big data and can introduce additional usability problems, often requiring data to be reformatted prior to use. This paper will explore how recent integrations between the GDR and the OEDI data lake have improved the accessibility and usability of geothermal data in a big way, making the data available to a broader audience, and enabling collaborative analysis and innovation across the greater geothermal industry.

access↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDR's success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

access↗

GOOML - Finding Optimization Opportunities for Geothermal Operations: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach. We have used this framework to develop digital twins that provide steamfield operators with an operational environment to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management for real world applications. The GOOML modeling software is built on a generic component-based systems framework that allows for both historical and forecast analysis. A GOOML model can perform historical data-assimilation using first-principal thermodynamics to create a meaningful data model. Historical production data can then be coupled with a forecast framework to train machine-learning models of steamfield components to predict future outputs. This modeling environment enables digital exploration of steamfield design configurations and operational scenarios. GOOML digital twins have been developed for steamfields in New Zealand and the United States representing differing power generation and field conditions. These digital twins have been validated by comparing hindcast predictions against historical production data. Reinforcement learning experiments were conducted to demonstrate the ability to programmatically explore the operations space using machine learning agents. Our initial results are compelling; two to five percent increases in annual energy production were demonstrated by the GOOML models with no additional infrastructure build required. GOOML offers a new approach to geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and interaction with digital twins. Through application of these tools, operators will realize greater availability and higher net generation which will increase the cost effectiveness of geothermal energy projects.

access↗

Improving the Accessibility and Usability of Geothermal Information with Data Lakes and Data Pipelines on the Geothermal Data Repository

The Geothermal Data Repository (GDR) provides universal access to data and information resulting from research and development activities funded by the Department of Energy (DOE). The GDR has extended this universal access to big data through integration with data lakes developed by the Open Energy Data Initiative (OEDI). Previously, large datasets such as seismic waveform or distributed acoustic sensing (DAS) data could only be accessed by institutions with high performance data storage and compute capabilities, effectively limiting the accessibility of big data to national labs, larger universities, and major corporations. Moreover, the time and resources needed to transport big data and configure them can produce additional barriers to use. Many of the standard formats used for structured data models (also known as content models) are incapable of handling big data and can introduce additional usability problems, often requiring data to be reformatted prior to use. This paper will explore how recent integrations between the GDR and the OEDI data lake have improved the accessibility and usability of geothermal data in a big way, making the data available to a broader audience, and enabling collaborative analysis and innovation across the greater geothermal industry.

access↗

Intelligent experiments through real-time AI: Fast Data Processing and Autonomous Detector Control for sPHENIX and future EIC detectors

This R&D project, initiated by the DOE Nuclear Physics AI-Machine Learning initiative in 2022, leverages AI to address data processing challenges in high-energy nuclear experiments (RHIC, LHC, and future EIC). Our focus is on developing a demonstrator for real-time processing of high-rate data streams from sPHENIX experiment tracking detectors. The limitations of a 15 kHz maximum trigger rate imposed by the calorimeters can be negated by intelligent use of streaming technology in the tracking system. The approach efficiently identifies low momentum rare heavy flavor events in high-rate p+p collisions (3MHz), using Graph Neural Network (GNN) and High Level Synthesis for Machine Learning (hls4ml). Success at sPHENIX promises immediate benefits, minimizing resources and accelerating the heavy-flavor measurements. The approach is transferable to other fields. For the EIC, we develop a DIS-electron tagger using Artificial Intelligence - Machine Learning (AI-ML) algorithms for real-time identification, showcasing the transformative potential of AI and FPGA technologies in high-energy nuclear and particle experiments real-time data processing pipelines.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

OmniXAS: A universal deep-learning framework for materials x-ray absorption spectra

X-ray absorption spectroscopy (XAS) is a powerful characterization technique for probing the local chemical environment of absorbing atoms. However, analyzing XAS data presents significant challenges, often requiring extensive, computationally intensive simulations, as well as significant domain expertise. These limitations hinder the development of fast, robust XAS analysis pipelines that are essential in high-throughput studies and for autonomous experimentation. Here, we address these challenges with OmniXAS, a framework that contains a suite of transfer learning approaches for XAS prediction, each uniquely contributing to improved accuracy and efficiency, as demonstrated on the K-edge spectra database covering eight 3⁢d transition metals (Ti–Cu). The OmniXAS framework is built upon three distinct strategies. First, we use M3GNet [Nat. Comput. Sci. 2, 718 (2022)] to derive latent representations of the local chemical environment of absorption sites as input for XAS prediction, achieving significant improvements over conventional featurization techniques. Second, we employ a hierarchical transfer learning strategy, training a universal multitask model across elements before fine-tuning for element-specific predictions. Models based on this cascaded approach after elementwise fine-tuning outperform element-specific models by up to 69%. Third, we implement cross-fidelity transfer learning, adapting a universal model to predict spectra generated by simulation of a different fidelity with a much higher computational cost. This approach improves prediction accuracy by up to 11% over models trained on the target fidelity alone. Our approach significantly boosts the throughput of XAS modeling by orders of magnitude as compared to first-principles simulations and is extendable to XAS prediction for a broader range of elements. The proposed transfer learning framework is generalizable to enhance deep-learning models that target other properties in materials research.

36 MATERIALS SCIENCE↗

CONSTRAINT-INDEPENDENT CONSTANT CTOA DETERMINATION FOR DUCTILE STABLE CRACK GROWTH

Crack tip opening angle (CTOA) has been used as a reliable fracture toughness parameter for decades to characterize stable ductile crack growth for thin-walled aerospace structures in the low-constraint conditions. Recently, the CTOA parameter was also applied to the pipeline industry, and a CTOA test standard ASTM E3039 was thus developed for testing a critical constant CTOA. Research showed that the constant CTOA can reasonably describe fracture toughness required to arrest a dynamic crack propagation for a modern gas pipeline. However, the CTOA fracture criterion requires constraint-independent CTOA toughness against stable ductile crack growth. ASTM E3039 recommends a drop weight tearing test (DWTT) specimen for CTOA testing. Since a shallow crack is used, DWTT measured CTOA may depend on constraint level at the crack tip. To understand if it is the case, this paper evaluates the critical CTOA for a set of fracture toughness tests on single edge notched bend (SENB) specimens with shallow and deep cracks based on four CTOA estimation models. In which, the Ln(P)-LLD linear fit model is similar to that used by ASTM E3039 in the CTOA calculation. Fracture test data for X80 pipeline steel and HY80 structural steel are considered in the CTOA evaluation. The results show that the four CTOA models can determine a crack size-independent constant CTOA over stable ductile crack growth for the SENB specimens. As a result, CTOA determined by ASTM E3039 is constraintindependent and transferable to use for an actual crack propagating in a gas pipeline.

Zhu, Xian-Kui↗