Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Investigating Building Energy Consumption and CO2 Emission in Phoenix Using AutoBEM and Future Typical Meteorological Year (fTMY) Weather Data

This research investigates the energy performance and CO2 emissions of each building stock across the Phoenix metropolitan area using the Automatic Building Energy Modeling (AutoBEM) framework and Model America v2 (MAv2) dataset from Oak Ridge National Laboratory (ORNL). Typical Meteorological Year (TMY) and Future Typical Meteorological Year (fTMY) files were used for AutoBEM simulation. The simulation results from TMY and fTMY were compared. It was found that a projected 10.28% increase in total CO2 emissions and a 9.30% rise in total energy consumption by 2080–2099 relative to current typical conditions. The results highlight the disparities in emissions among different building stocks and the influence of climate change on future energy demand. The findings underscore the necessity of targeted policy interventions and retrofitting strategies (eg. advanced HVAC systems, improved insulation, reflective roofing) to mitigate emissions in high-energy-use and emission-intensed buildings, particularly as climate conditions evolve. This study contributes to the growing understanding of building-sector emissions and their long-term implications under future climate scenarios.

Li, Hang [ORNL] (ORCID:0000000306001920)↗

Investigating Scientific Data Change with User Research Methods

Scientific datasets are continually expanding and changing due to fluctuations with instruments, quality assessment and quality control processes, and modifications to software pipelines. Datasets include minimal information about these changes or their effects requiring scientists manually assess modifications through a number of labor intensive and ad-hoc steps. The Deduce project is investigating data change to develop metrics, methods, and tools that will help scientists systematically identify and make decisions around data changes. Currently, there is a lack of understanding, and common practices, for identifying and evaluating changes in datasets since systematically measuring and managing data change is under explored in scientific work. We are conducting user research to address this need by exploring scientist's conceptualizations, behaviors, needs, and motivations when dealing with changing datasets. Our user research utilizes multiple methods to produce foundational, generative insights and evaluate research products produced by our team. In this paper, we detail our user research process and outline our findings about data change that emerge from our studies. Our work illustrates how scientific software teams can push beyond just usability testing user interfaces or tools to better probe the underlying ideas they are developing solutions to address.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Artificial Intelligence for Data Center Operations (AIOps): Cooperative Research and Development (Final Report)

High performance computing data centers will increasingly need to rely on automation to keep pace with exascale growth in compute capability and to manage and optimize the data center environment and facility resources. Artificial intelligence and machine learning approaches provide the means to improve HPC data center operational efficiency, by learning historical trends and training models to operate on real-time data collected from both IT and facilities sources. NREL has developed methods of real-time collection, aggregation and streaming of these data in the ESIF HPC Data Center and has collected a significant dataset of relevant metrics across computer systems, racks, environmental, building and utility sources for research into various predictive analytics problems. HPE's Advanced Technology Group (ATG) is doing comprehensive research into exascale monitoring and management for High Performance Computing (HPC) systems (hereinafter HPE's Data Monitoring/ Management Technology). NREL and HPE will collaborate to add Artificial Intelligence (AI) to NREL's real-time data collection/ aggregation/ streaming system and HPE's Data Monitoring/ Management System, with the goal of improving the operational efficiency of NREL's Energy Systems Integration Facility (ESIF) HPC Data Center through data analytics on both historical and real-time data from IT systems and facilities operations. This collaboration will consist of efforts in Data Management, Data Analytics, and AI/ML Optimization for both manual and autonomous intervention in data center operations. This will be a multi-year, multi-staged effort with a goal towards building capabilities for an Advanced Smart Facility, and demonstration of these techniques in the NREL ESIF HPC Data Center.

97 MATHEMATICS AND COMPUTING↗

Experiences with a Flexible User Research Process to Build Data Change Tools

Scientific software development processes are understood to be distinct from commercial software development practices due to uncertain and evolving states of scientific knowledge. Sustaining these software products is a recognized challenge, but under-examined is the usability and usefulness of such tools to their scientific end users. User research is a well-established set of techniques (e.g., interviews, mockups, usability tests) applied in commercial software projects to develop foundational, generative, and evaluative insights about products and the people who use them. Currently these approaches are not commonly applied and discussed in scientific software development work. The use of user research techniques in scientific environments can be challenging due to the nascent, fluid problem spaces of scientific work, varying scope of projects and their user communities, and funding/economic constraints on projects.In this paper, we reflect on our experiences undertaking a multi-method user research process in the Deduce project. The Deduce project is investigating data change to develop metrics, methods, and tools that will help scientists make decisions around data change. There is a lack of common terminology since the concept of systematically measuring and managing data change is under explored in scientific environments. To bridge this gap we conducted user research that focuses on user practices, needs, and motivations to help us design and develop metrics and tools for data change. This paper contributes reflections and the lessons we have learned from our experiences. We offer key takeaways for scientific software project teams to effectively and flexibly incorporate similar processes into their projects.

97 MATHEMATICS AND COMPUTING↗

Size-resolved Eddy-Covariance Particle Flux Measurement during the TRACER Campaign (Final Report)

The main goal of the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign was to study aerosol–cloud interactions during deep convection over the Houston area. This project deployed a suite of instrumentation with the aim to (1) quantify turbulent vertical particle fluxes during at DOE-ARM sites, including TRACER, (2) assess hygroscopic growth factors and hygroscopicity parameters of the material driving modal aerosol growth during new particle formation and growth events, (3) derive turbulent aerosol mass fluxes using co-located Doppler LIDAR measurements, and (4) create quality-controlled PI data products to support future research utilizing data collected during the TRACER campaign. This report summarized the main findings from the deployments at two DOE-ARM sites. Briefly, we found that new particle formation may occur aloft, in a residual layer, near the top of the boundary layer. Small grown particles appear later due to downward mixing with daytime turbulence. The species that are responsible for aerosol modal growth had hygroscopicity parameters varying between 0.05 and 0.34. These values systematically depended on the wind sector, suggesting that the chemical composition of the precursors differed. This work demonstrated that lidar retrievals of the elastic backscatter and Doppler velocity can be used to obtain surface number emissions of particles with a diameter greater than 0.53 µm. During TRACER, emission particle number fluxes peaked near ∼ 100 cm−2 s−1. Multiple quality-controlled PI data products that will support future TRACER related science were generated and made publically available.

54 ENVIRONMENTAL SCIENCES↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.↗

From Simple Labels to Time-Use Integrations: Supporting the Spectrum of Qualitative Travel Behavior Data

In transportation research, applications of travel behavior data collection are context-specific and require different types of qualitative inputs. These inputs can be viewed as spanning a spectrum of user burden and data quality, from simple trip labels to complex time-use surveys. However, each currently active smartphone-based travel diary platform appears to only support one type of qualitative input, and the effort required for customization is unclear. In this paper, we characterize the spectrum by defining four canonical use cases: (i) trip labels, (ii) trip questionnaire, (iii) counterfactual trips, and (iv) time-use surveys. We then outline a mechanism for supporting configurable user inputs on the same underlying smartphone-based sensing mechanism and demonstrate that it can support all the use cases without any code changes. We further demonstrate that the flexible data model that underpins this mechanism can enable real-time monitoring and analysis. Finally, we evaluate per-user data collection and engagement metrics for large-scale deployments of three canonical use cases, spanning 10 programs, 435 users, and 251,041 trips, and a maximum duration of 800 days. Future efforts may support additional use cases through an expanded configuration and provide greater insight into user engagement. We hope that these insights enable the research community to look at qualitative inputs through a new lens and experiment with novel use cases to fill in the spectrum.

ADVANCED PROPULSION SYSTEMS,POWER TRANSMISSION AND↗

Expanding the access of wearable silicone wristbands in community-engaged research through best practices in data analysis and integration

Wearable silicone wristbands are a rapidly growing exposure assessment technology that offer researchers the ability to study previously inaccessible cohorts and have the potential to provide a more comprehensive picture of chemical exposure within diverse communities. However, there are no established best practices for analyzing the data within a study or across multiple studies, thereby limiting impact and access of these data for larger meta-analyses. We utilize data from three studies, from over 600 wristbands worn by participants in New York City and Eugene, Oregon, to present a first-of-its-kind manuscript detailing wristband data properties. We further discuss and provide concrete examples of key areas and considerations in common statistical modeling methods where best practices must be established to enable meta-analyses and integration of data from multiple studies. Finally, we detail important and challenging aspects of machine learning, meta-analysis, and data integration that researchers will face in order to extend beyond the limited scope of individual studies focused on specific populations.

Bramer, Lisa M.↗

QUOTAS: A New Research Platform for the Data-driven Discovery of Black Holes

We present QUOTAS, a novel research platform for the data-driven investigation of supermassive black hole (SMBH) populations. While SMBH data—observations and simulations—have grown in complexity and abundance, our computational environments and tools have not matured commensurately to exhaust opportunities for discovery. To explore the BH, host galaxy, and parent dark matter halo connection—in this pilot version—we assemble and colocate the high-redshift, z > 3 quasar population alongside simulated data at the same cosmic epochs. As a first demonstration of the utility of QUOTAS, we investigate correlations between observed Sloan Digital Sky Survey (SDSS) quasars and their hosts with those derived from simulations. Leveraging machine-learning algorithms (ML), to expand simulation volumes, we show that halo properties extracted from smaller dark-matter-only simulation boxes successfully replicate halo populations in larger boxes. Next, using the Illustris-TNG300 simulation that includes baryonic physics as the training set, we populate the larger LEGACY Expanse dark-matter-only box with quasars, and show that observed SDSS quasar occupation statistics are accurately replicated. First science results from QUOTAS comparing colocated observational and ML-trained simulated data at z3 are presented. QUOTAS demonstrates the power of ML, in analyzing and exploring large data sets, while also offering a unique opportunity to interrogate theoretical assumptions that underpin accretion and feedback models. QUOTAS and all related materials are publicly available at the Google Kaggle platform. (The full data set—observational data and simulation data—are available at: https://www.kaggle.com/ and the codes are available at:https://www.kaggle.com/datasets/quotasplatform/quotas)

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

MISIP: a data standard for the reuse and reproducibility of any stable isotope probing-derived nucleic acid sequence and experiment

DNA/RNA-stable isotope probing (SIP) is a powerful tool to link in situ microbial activity to sequencing data. Every SIP dataset captures distinct information about microbial community metabolism, process rates, and population dynamics, offering valuable insights for a wide range of research questions. Data reuse maximizes the information derived from the labor and resource-intensive SIP approaches. Yet, a review of publicly available SIP sequencing metadata showed that critical information necessary for reproducibility and reuse was often missing. Here, we outline the Minimum Information for any Stable Isotope Probing Sequence (MISIP) according to the Minimum Information for any (x) Sequence (MIxS) framework and include examples of MISIP reporting for common SIP experiments. Our objectives are to expand the capacity of MIxS to accommodate SIP-specific metadata and guide SIP users in metadata collection when planning and reporting an experiment. The MISIP standard requires 5 metadata fields—isotope, isotopolog, isotopolog label, labeling approach, and gradient position—and recommends several fields that represent best practices in acquiring and reporting SIP sequencing data (e.g., gradient density and nucleic acid amount). The standard is intended to be used in concert with other MIxS checklists to comprehensively describe the origin of sequence data, such as for marker genes (MISIP-MIMARKS) or metagenomes (MISIP-MIMS), in combination with metadata required by an environmental extension (e.g., soil). The adoption of the proposed data standard will improve the reuse of any sequence derived from a SIP experiment and, by extension, deepen understanding of in situ biogeochemical processes and microbial ecology.

Simpson, Abigayle↗

Artificial Intelligence in Nuclear Physics

Artificial Intelligence (AI) and Machine Learning (ML) are rapidly developing fields providing data-driven algorithms to predict, classify, and make decisions based on data. Nuclear Physics Research is data-driven and AI/ML techniques have been implemented for experiment and accelerator control, in theoretical applications, and in data processing and analysis. These algorithms open possibilities for automation, thereby augmenting human capabilities. Additionally, Open Science is enabled by simultaneous analyses of multiple data sources, leading to scientific knowledge. This talk will summarize current applications of AI/ML in nuclear physics, as well as accelerator applications, and will cover upcoming initiatives and research in AI/ML.

Jeske, Torri↗

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

Natural Language Processing-Enhanced Nuclear Industry Operating Experience Data Analysis to Support Risk Model Parameter Estimations

This set of slides has been prepared for a talk at the INL AI/ML Symposium held on September 8, 2022. Presentation outline: Background - Nuclear power plant operating experience data sources • Research focus and motivation - Analyzing free-text operating experience data: present and future • Research method - Input - Methodological steps - Output • Conclusions and next steps

99 GENERAL AND MISCELLANEOUS↗

C-HER Metadata Overview: Approach, Standards, and Rigor for the Centralized Health and Exposomic Resource

The Centralized Health and Exposomic Resource (C-HER) unifies environmental, demographic, geographic, and health-related data for exposomic research. The source data differ in format, geographic coverage, time period, resolution, terminology, and documentation. We use a common metadata framework to describe those differences and to record how each data resource has been processed, documented, and ingested. This document relates only to the C-HER metadata framework. It explains the information that is recorded for each resource, the standards used to organize that information, the conditions for metadata completeness, and the relationship between metadata and quality review. It is intended for those who need to understand what C-HER metadata communicates and how it supports appropriate use of the data. It is not an implementation specification or procedure. It does not document the database schema, source code, deployment configuration, transformation algorithms, or dataset-specific QA/QC thresholds. Those materials are maintained separately.

MacFarland, Midgie [ORNL] (ORCID:0009000807354078)↗

Scalable GPS Data Logging To Support Advanced Fleet Analysis

This highlight details the key takeaways from a project that utilized NLR's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline. National Laboratory of the Rockies researchers developed and demonstrated low-cost, open-source Arduino data loggers with 3D-printed cases that are compatible with global navigational systems and built with components available ubiquitously worldwide, enabling cost-effective collection and analysis of fleet operational data. Validated on an overseas transit bus fleet, NLR analysis showed that, with sufficient charging opportunities, 90% of observed duty cycles could be accomplished by electric buses with no modifications to operations.

33 ADVANCED PROPULSION SYSTEMS↗

U.S. Manufacturing Water Use Data and Estimates: Current State, Limitations, and Future Needs for Supporting Manufacturing Research and Development

Water is essential to manufacturing operations; without it, many facilities could not operate or meet production demands. Physical, reputational, and regulatory risks to water supplies compounded by climate change-induced impacts on hydrological conditions threaten the adequacy of water supplies for manufacturing. Manufacturing water use has not been a major focus of either water or manufacturing-related research. Research and development (R&D) aimed at helping manufacturers use water more sustainably and adapt to changing water conditions is needed to ensure a thriving sector and economy. However, the ability to identify R&D needs is severely limited due to a lack of current, statistically representative data on manufacturing water use and its environmental implications. In this Perspective, we outline four key questions to inform R&D on manufacturing use and highlight how the current state of water data in the United States does not support the adequate investigation of these questions. We make recommendations for the water data characteristics needed to explore the research questions and knowledgeably inform R&D on manufacturing water use.

McCall, James↗

The Zooplankton International Geospatial (ZIG) dataset: A global repository of spatiotemporal freshwater zooplankton community composition data to support ecological research

Zooplankton play critical roles in aquatic ecosystem function and food webs. Nevertheless, global syntheses of their abundance and community dynamics are challenging due to methodological differences across monitoring programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we assembled, curated, validated, and harmonized the Zooplankton International Geospatial (ZIG) dataset, which includes co-located and contemporaneous zooplankton, water chemistry, and limnological data from 307 lakes and reservoirs. ZIG includes waterbodies from each major lake thermal region and range in size from 0.8-2,805,8600 hectares. Temporal coverage for individual waterbodies ranges between 1-60 years of data (median = 4 years) with sampling from once annually to weekly. ZIG is publicly available and can be used to understand freshwater biodiversity change and its drivers at unprecedented scales, and we consider it to be a cornerstone for future investigations of freshwater biology, chemistry, and ecology.

Figary, Stephanie [Cornell University, Ithaca, NY]↗