Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Massive all-atom analysis of 2D materials with quantum properties (Final report)

Improvements in microscopy have enabled the acquisition of data at a scale that is difficult to process manually, making automated machine learning approaches to analyzing experimental images essential. In this project, we developed and applied machine learning (ML) workflows for atomic resolution scanning transmission electron microscopy (STEM) images. This development included improving both methodology as well as generating user-friendly codes. We developed machine learning architectures which, after training, automatically identify the location and types of defects throughout a material. We used these data to produce class-averaged images of 2D atomic coordinates with up to 0.3 pm precision, uncovering the structure and oscillations of long-range strain fields around point defects in WSe 2-2x Te 2x . We also resolved a long-standing problem in this field in the training of ML models, a lack of labeled experimental data, by developing a cycle-GAN that transformed simulated-generated labeled data into labeled data indistinguishable from experiment and therefore suitable for training. This removed the remaining parts of the ML data processing workflow where human intervention was still critical and therefore a bottleneck to working at scale. Codes have been developed and released for this full machine learning workflow. ML approaches to partially automate STEM acquisition were also developed. Finally we applied ML and other advanced data processing methods to several materials science problems in two-dimensional materials, including studying the evolution of hyperuniformity with defect concentration in WSe2, understanding phase transformations in transition metal dichalcogenides during in-situ heating in the STEM, and exploring how 2D interfaces transform from twisted into aligned structures.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Plant science decadal vision 2020–2030: Reimagining the potential of plants for a healthy and sustainable future

Abstract Plants, and the biological systems around them, are key to the future health of the planet and its inhabitants. The Plant Science Decadal Vision 2020–2030 frames our ability to perform vital and far‐reaching research in plant systems sciences, essential to how we value participants and apply emerging technologies. We outline a comprehensive vision for addressing some of our most pressing global problems through discovery, practical applications, and education. The Decadal Vision was developed by the participants at the Plant Summit 2019, a community event organized by the Plant Science Research Network. The Decadal Vision describes a holistic vision for the next decade of plant science that blends recommendations for research, people, and technology. Going beyond discoveries and applications, we, the plant science community, must implement bold, innovative changes to research cultures and training paradigms in this era of automation, virtualization, and the looming shadow of climate change. Our vision and hopes for the next decade are encapsulated in the phrase reimagining the potential of plants for a healthy and sustainable future. The Decadal Vision recognizes the vital intersection of human and scientific elements and demands an integrated implementation of strategies for research (Goals 1–4), people (Goals 5 and 6), and technology (Goals 7 and 8). This report is intended to help inspire and guide the research community, scientific societies, federal funding agencies, private philanthropies, corporations, educators, entrepreneurs, and early career researchers over the next 10 years. The research encompass experimental and computational approaches to understanding and predicting ecosystem behavior; novel production systems for food, feed, and fiber with greater crop diversity, efficiency, productivity, and resilience that improve ecosystem health; approaches to realize the potential for advances in nutrition, discovery and engineering of plant‐based medicines, and "green infrastructure." Launching the Transparent Plant will use experimental and computational approaches to break down the phytobiome into a "parts store" that supports tinkering and supports query, prediction, and rapid‐response problem solving. Equity, diversity, and inclusion are indispensable cornerstones of realizing our vision. We make recommendations around funding and systems that support customized professional development. Plant systems are frequently taken for granted therefore we make recommendations to improve plant awareness and community science programs to increase understanding of scientific research. We prioritize emerging technologies, focusing on non‐invasive imaging, sensors, and plug‐and‐play portable lab technologies, coupled with enabling computational advances. Plant systems science will benefit from data management and future advances in automation, machine learning, natural language processing, and artificial intelligence‐assisted data integration, pattern identification, and decision making. Implementation of this vision will transform plant systems science and ripple outwards through society and across the globe. Beyond deepening our biological understanding, we envision entirely new applications. We further anticipate a wave of diversification of plant systems practitioners while stimulating community engagement, underpinning increasing entrepreneurship. This surge of engagement and knowledge will help satisfy and stoke people's natural curiosity about the future, and their desire to prepare for it, as they seek fuller information about food, health, climate and ecological systems.

59 BASIC BIOLOGICAL SCIENCES↗

Automating the interpretation of PM 2.5 time–resolved measurements using a data–driven approach

The rapid development of automated measurement equipment enables researchers to collect greater quantities of time-resolved data from indoor and outdoor environments. While significant, the interpretation of the resulting data can be a time-consuming effort. This paper introduces an automated process of interpreting PM 2.5 time-resolved data and differentiating PM 2.5 emissions resulting from indoor and outdoor sources. Here, we use Random Forest (RF), a machine learning approach, to study a dataset of 836 indoor emission events that occurred over a 2-week period in 18 apartments in California. In this paper, we show model development and evaluate its performance as the sample size and source vary. We discuss the characteristics of the dataset that tended to help the source identification and why. For example, we show that data from many events and from different apartments are essential for the model to be suitable for analyzing a new separate dataset. We also show that longitudinal data appear to be more helpful than the time frequency of measurements within a given apartment. We use the resulting RF model to analyze PM 2.5 data of an entirely separate dataset collected from 65 new homes in California. The RF model identifies 442 indoor emission events, with only a few misidentifications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

WFIP3 - BLOC site - CU Profiling Lidar (Windcube v1) / Raw data

This dataset is from WFIP3 BLOC site WindCube v1 Profiling Lidar / Raw Data from CU Boulder. Please note that the .sta files are in human-readable text, while the .rtd files require some processing software that is not easily automated. We plan to process the .rtd data following the campaign, but if you wish to try it sooner, we can make it available upon request.

17 WIND ENERGY↗

WFIP3 - BLOC site - CU Profiling Lidar (Windcube v1) / Raw data

This dataset is from WFIP3 BLOC site WindCube v1 Profiling Lidar / Raw Data from CU Boulder. Please note that the .sta files are in human-readable text, while the .rtd files require some processing software that is not easily automated. We plan to process the .rtd data following the campaign, but if you wish to try it sooner, we can make it available upon request.

17 WIND ENERGY↗

WFIP3 - BLOC site - CU Profiling Lidar (Windcube v1) / Raw data

This dataset is from WFIP3 BLOC site WindCube v1 Profiling Lidar / Raw Data from CU Boulder. Please note that the .sta files are in human-readable text, while the .rtd files require some processing software that is not easily automated. We plan to process the .rtd data following the campaign, but if you wish to try it sooner, we can make it available upon request.

17 WIND ENERGY↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Residential Building Energy Efficiency Field Studies: Low-Rise Multifamily

In recent years, the U.S. Department of Energy (DOE) has conducted a series of research studies to validate energy efficient building technologies in the field. Much of the work has focused on single-family construction, and some has also addressed commercial energy codes. The work detailed in this DOE-funded study (EE0007616) focuses on low-rise multifamily buildings (three stories or fewer above grade) in various regions of the United States, and reports on how state-level building codes are being implemented, both in terms of observed characteristics and also in terms of estimated energy impacts. Nearly 100 buildings across four states—Illinois, Minnesota, Oregon, and Washington—were sampled, which represent a range of climate types from mild temperature to very cold continental. Both common entry and outdoor entry buildings were included, and a parallel research project evaluated envelope air tightness and current still-evolving air tightness testing methods. Finally, a set of structured interviews of building designers and other relevant professionals was carried to out to gain more insight into this market. To the greatest extent possible, the methodology developed under the project for low-rise multifamily buildings mirrored the approach established by Pacific Northwest National Laboratory (PNNL) for single-family residential buildings (https://www.energy.gov/eere/buildings/downloads/residential-building-energy-code-field-study). This included the general approach to sampling, recruitment, and data collection, as well as data analysis and presentation. The range of permitting dates for the sites encompassed two energy code cycles in most regions. All states in the study had adopted a variation of the International Energy Conservation Code (IECC) for the structure of their state code. The low-rise multifamily occupancy presents a hybrid building type: most of the building’s conditioned floor area was covered by the residential chapter of the code while portions of the building (such as corridors and common spaces) fell under the commercial code chapter. The key items assessed in this work were: Building Shell—exterior wall insulation, ceiling insulation, foundation insulation, windows. Common Areas—HVAC and lighting. Living Units—lighting, ventilation. A few items were not assessed in detail, given their relative paucity in this occupancy type; these included duct leakage, pipe insulation, and hot water circulation controls. Building characteristics were collected via a combination of architectural, mechanical, electrical, and plumbing plan reviews and field inspections, and entered into a spreadsheet-based tool that was later queried to build a database. Data went through quality control both upon arrival and via a later semi-automated review and assurance process. Most of the data are presented graphically so that the reader can quickly assess compliance with the applicable energy codes (both by state and by code year). As a final step, EnergyPlus™ simulations were created for all buildings in the study to estimate both the as-found energy use intensity (EUI) and the energy and CO 2 that could be saved if features that were found to not meet code minimums were brought up to code. The savings estimates were tabulated for each of the four states in the study. The research team found that the single-family approach was largely applicable to low-rise multifamily buildings. This applies to both the data collection and the prototype EUI analysis. Most of the occupied space is living units and falls under residential energy codes, and many characteristics use similar envelope construction and relatively straightforward mechanical systems and lighting. One of the most challenging aspects of this work was to build an effective spreadsheet-based data collection instrument that could allow efficient collection of both building plan and field data. The research team is of the view that other methods could be equally effective if the work is done carefully with diligent quality control. The primary findings for the work center around the thermal envelope and mechanical systems and lighting at the sites: For thermal envelope components, the majority of buildings met or were better than the prescriptive code.This suggests that building designers and builders are aware of code requirements. In some cases, surveyed buildings were designed to qualify for energy efficiency certification programs. These buildings made up at least 20% of sampled buildings in each state. Almost all buildings met mechanical system efficiency requirements (for both living units and common areas). In some cases, sites employed systems that were considerably more efficient than required by the applicable energy code. Dwelling units had a majority of high-efficacy lighting, often in excess of the state’s residential code requirements. While high-efficacy fixtures were also typical in common areas (corridors and stairwells), lighting power densities (LPDs) in these areas were sometimes higher than levels dictated by the applicable part of the state commercial energy code. The simulation models run on a series of low-rise multifamily prototypes, informed by a composite of the field data collected, calculated annual EUIs of between 20 and 50 kBtu/ft2-yr, with the range representing the effects of both building characteristics and building location (climate zone). A detailed process (based on simulations of prototype buildings) was used to estimate the amount of avoided energy use that would occur if 100% adherence to energy codes were attained. The results indicated modest savings are attainable for items such as window thermal performance and common area lighting. The result is overall only a modest potential for additional energy savings, averaging about 10% of EUI.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

In situ Synchrotron X‐ray Metrology Boosted by Automated Data Analysis for Real‐time Monitoring of Cathode Calcination

Abstract Synchrotron X‐ray‐based in situ metrology is advantageous for monitoring the synthesis of battery materials, offering high throughput, high spatial and temporal resolution, and chemical sensitivity. However, the rapid generation of massive data poses a challenge to on‐site, on‐the‐fly analysis needed for real‐time process monitoring. Here, a weighted lagged cross‐correlation (WLCC) similarity approach is presented for automated data analysis, which merges with in situ synchrotron X‐ray diffraction metrology to monitor the calcination process of the archetypal nickel‐based cathode, LiNiO 2 . The WLCC approach, incorporating variables that account for peak shifts and width changes associated with structural transformations, enables rapid extraction of phase progression within 10 seconds from tens of diffraction patterns. Details are captured, from initial precursors to intermediates and the final layered LiNiO 2 , providing information for agile on‐site adjustments during experiments and complementing post hoc diffraction analysis by offering insights into early‐stage phase nucleation and growth. Expanding this data‐powered platform paves the way for real time calcination process monitoring and control, which is pivotal to quality control in battery cathode manufacturing.

36 MATERIALS SCIENCE↗

Advanced Tritium Process Analytics and Optimization via a Digital Twin, SRNL-TR-2023-00550

Improving the process knowledge and understanding of TCAP can be achieved by incorporating advanced tools. This project addresses “Advanced analytics for modeling, forecasting, & optimization of the tritium refinement process” and is being applied to the TCAP process. The value of digitization and advanced analytics for TCAP data are tracking of gas mixtures & inventories is presently a manual effort that is rather cumbersome. Incorporating some automation into the data capture will reduce errors and storing data in an accessible database that can be used for tracking as well as process improvements. In addition, the database approach will enable longer term history to be maintained rather than the current practice of deleting data after six months. The digitized and automated data will allow for models to be developed for the process and will enable the forecasting and optimization.

Korinko, Paul S.↗

Precision Plant Biomass Characterization in Agriculture: Harnessing Machine Learning and Hyperspectral Imaging [Slides]

Efficient Biomass Separation Object detection of anatomical parts (Cob, Stalk, Husk) in IR images enables precise separation, improving preprocessing (e.g., drying, grinding) for biofuel production. Detailed Biomass Characterization with Hyperspectral Data Hyperspectral imaging captures spectral signatures of biomass, allowing for the identification of specific traits like moisture content, lignin levels, and nutrient composition, leading to optimized treatments for each biomass part. Enhanced Feedstock Quality By leveraging hyperspectral data, feedstock can be processed based on its chemical composition, improving conversion efficiency and biofuel yield. Automation for Large-Scale Operations Automated object detection and hyperspectral data analysis reduce manual labor, ensuring accurate sorting and faster processing, making large-scale biofuel production more efficient. Maximized Biomass Utilization Accurate identification of biomass properties minimizes waste and ensures that each part is processed according to its highest biofuel potential.

09 BIOMASS FUELS↗

Evaluation of Converter Performance Considering Static and Dynamic Device Part-to-Part Variability

This paper presents a methodology to incorporate and analyze the impact of semiconductor device part-to-part variation on power converter performance. By integrating extensive static and dynamic device characterization data with an automated compact model generation process that reflects manufacturing variability, device models with inherent variability features are utilized in converter simulations for a comprehensive assessment of performance impacts. The traditional converter performance evaluation process typically yields fixed efficiency values, often dismissing the inherent part-to-part variability caused by the manufacturing process of semiconductor devices. To address this limitation, a large population of devices was characterized to capture variations in static parameters-such as transfer, output, and capacitance characteristics-as well as dynamic behaviors, including switching losses. This data-driven approach enables the development of individual compact models, which were then integrated into converter simulations to evaluate efficiency ranges rather than single point estimated values. The converter simulation results show that part-to-part component variation can lead to significant efficiency deviations, exceeding several percentage points in high-power conversion applications. By offering a more accurate representation of converter behavior under real-world manufacturing conditions, this methodology enables designers to anticipate performance variability, improving the robustness of power converter designs.

device characterization↗

Eukaryotic genomes from a global metagenomic data set illuminate trophic modes and biogeography of ocean plankton

ABSTRACT Metagenomics is a powerful method for interpreting the ecological roles and physiological capabilities of mixed microbial communities. Yet, many tools for processing metagenomic data are neither designed to consider eukaryotes nor are they built for an increasing amount of sequence data. EukHeist is an automated pipeline to retrieve eukaryotic and prokaryotic metagenome-assembled genomes (MAGs) from large-scale metagenomic sequence data sets. We developed the EukHeist workflow to specifically process large amounts of both metagenomic and/or metatranscriptomic sequence data in an automated and reproducible fashion. Here, we applied EukHeist to the large-size fraction data (0.8–2,000 µm) from Tara Oceans to recover both eukaryotic and prokaryotic MAGs, which we refer to as TOPAZ (Tara Oceans Particle-Associated MAGs). The TOPAZ MAGs consisted of >900 environmentally relevant eukaryotic MAGs and >4,000 bacterial and archaeal MAGs. The bacterial and archaeal TOPAZ MAGs expand upon the phylogenetic diversity of likely particle- and host-associated taxa. We use these MAGs to demonstrate an approach to infer the putative trophic mode of the recovered eukaryotic MAGs. We also identify ecological cohorts of co-occurring MAGs, which are driven by specific environmental factors and putative host-microbe associations. These data together add to a number of growing resources of environmentally relevant eukaryotic genomic information. Complementary and expanded databases of MAGs, such as those provided through scalable pipelines like EukHeist, stand to advance our understanding of eukaryotic diversity through increased coverage of genomic representatives across the tree of life. IMPORTANCE Single-celled eukaryotes play ecologically significant roles in the marine environment, yet fundamental questions about their biodiversity, ecological function, and interactions remain. Environmental sequencing enables researchers to document naturally occurring protistan communities, without culturing bias, yet metagenomic and metatranscriptomic sequencing approaches cannot separate individual species from communities. To more completely capture the genomic content of mixed protistan populations, we can create bins of sequences that represent the same organism (metagenome-assembled genomes [MAGs]). We developed the EukHeist pipeline, which automates the binning of population-level eukaryotic and prokaryotic genomes from metagenomic reads. We show exciting insight into what protistan communities are present and their trophic roles in the ocean. Scalable computational tools, like EukHeist, may accelerate the identification of meaningful genetic signatures from large data sets and complement researchers’ efforts to leverage MAG databases for addressing ecological questions, resolving evolutionary relationships, and discovering potentially novel biodiversity.

59 BASIC BIOLOGICAL SCIENCES↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Artificial Intelligence in Nuclear Physics

Artificial Intelligence (AI) and Machine Learning (ML) are rapidly developing fields providing data-driven algorithms to predict, classify, and make decisions based on data. Nuclear Physics Research is data-driven and AI/ML techniques have been implemented for experiment and accelerator control, in theoretical applications, and in data processing and analysis. These algorithms open possibilities for automation, thereby augmenting human capabilities. Additionally, Open Science is enabled by simultaneous analyses of multiple data sources, leading to scientific knowledge. This talk will summarize current applications of AI/ML in nuclear physics, as well as accelerator applications, and will cover upcoming initiatives and research in AI/ML.

Jeske, Torri↗