Engineering PapersSearch

SEARCH · Engineering Papers

Results for “CAN data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Long‐Term Large‐Scale Atmospheric Forcing Data From Three‐Dimensional Constrained Variational Analysis for the ARM SGP Site

Here, this study presents a long‐term three‐dimensional large‐scale forcing data set (VARANAL3D) derived from the three‐dimensional constrained variational analysis (3DCVA) method at the Atmospheric Radiation Measurement (ARM) program Southern Great Plains (SGP) site from 2004 to 2018. Building on the same input data sets as the conventional continuous forcing data set (VARANAL), VARANAL3D maintains overall consistency in domain‐averaged fields while introducing spatial variability, offering critical insights into the influence of mesoscale synoptic systems on cloud‐related processes. Evaluations are conducted across four cloud and precipitation regimes: Clear‐sky, Shallow‐clouds, Afternoon‐precipitation, and Nocturnal‐precipitation, presenting high consistency of the domain‐mean forcing data sets while emphasizing the role of subdomain forcing variability particularly in precipitating regimes. Single column model (SCM) simulations demonstrate that subdomain VARANAL3D forcing improves cloud and precipitation representation, with the ensemble outperforming domain‐mean forcing in three cloudy and precipitating regimes. Overall, these results highlight VARANAL3D's value for investigating the impacts of spatial variability of large‐scale forcing on atmospheric processes. The VARANAL3D data set provides new opportunities for evaluating model physics, advancing the development of scale‐aware parameterizations and deepening our understanding of cloud and precipitation dynamics.

Environmental sciences

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility

Data-Conforming Data-Driven Control: Avoiding Premature Generalizations Beyond Data

Data-driven and adaptive control approaches face the problem of introducing sudden distributional shifts beyond the distribution of data encountered during learning. Therefore, they are prone to invalidating the very assumptions used in their own construction. This is due to the linearity of the underlying system, inherently assumed and formulated in most data-driven control approaches, which may falsely generalize the behavior of the system beyond the behavior experienced in the data. This article seeks to mitigate these problems by enforcing consistency of the newly designed closed-loop systems with data and slowing down any distributional shifts in the joint state-input space. This is achieved through incorporating affine regularization terms and linear matrix inequality constraints to data-driven approaches, resulting in convex semi-definite programs that can be efficiently solved by standard software packages. We discuss the optimality conditions of these programs and then conclude this article with a numerical example that further highlights the problem of premature generalization beyond data and shows the effectiveness of our proposed approaches in enhancing the safety of data-driven control methods.

97 MATHEMATICS AND COMPUTING

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES

Terrestrial laser scanning data (Levels 0 and 1) from Urban Biogeochemistry Pilot Project sites, Knoxville, Tennessee, Jul 2024 - Jul 2025

This data package contains data from terrestrial laser scanning (TLS) at five urban park sites in Knoxville, Tennessee, USA. All parks include open-grown and/or closed-canopy trees and mixed nearby land use. These study sites were established as part of the Urban Biogeochemistry Pilot Project, which has an overall goal of better understanding how hydrobiogeochemical cycling is altered within the human environment. These five sites represent a gradient of urbanization, and were instrumented to understand hydrological and biogeochemical cycling (e.g., soil moisture, soil physical properties and biogeochemistry, tree transpiration, species type). The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree- and stand-level characterization of woody structure and leaf area. TLS scans were placed to capture the area around trees with sap flow sensors, and as much of a 50 m radius area around the meteorological station as possible given site property limits. Derived products will allow upscaling of water content and transpiration data. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection

DOE EV Data Collection - Vehicle Data

Vehicle data consist of electric vehicle performance data collected directly from the vehicle during standard operations. Data were collected using onboard data loggers that were either installed by the project team or preinstalled by the original equipment manufacturer. Data recorded by the data loggers were made accessible via an online web portal or an application programming interface. Different data loggers were used (HEM, ViriCiti, and Geotab), and the method for each vehicle is defined in the vehicle attributes file. Some systems collected data on a “trip-level” basis, in which each row of a table represents a single trip (the period between a key-on and key-off event), whereas other data were collected on a per-day basis, in which each row represents a single day of operation. Data were collected over a range of data collection periods, depending on the project. Data have been anonymized by removing information or decreasing information resolution as necessary so that fleets are not identifiable. Due to the wide range of vehicle types represented and variation in data collection, data parameters and frequencies differ between vehicles and fleets The **Performance Data Daily/Trip Data Dictionaries** contain definitions for each available parameter associated with a vehicle’s operations, aggregated at either a daily or trip level. The parameters available will vary from vehicle to vehicle, but every possible parameter will be defined. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. The **Vehicle Data** tables contain the data from each vehicle’s operations, aggregated at either a daily or trip level, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

DOE EV Data Collection - Charging Data

Charging data are collected from one of three sources, each with varying levels of additional information. These sources, in approximate order from most to least additional information, are: • The electric vehicle supply equipment (charger) • Onboard the vehicle itself • From a utility submeter. Many chargers provide software that allows for the collection and reporting of charging session data. If unavailable, data may be recorded by the charging vehicle’s onboard systems. If neither of these options is available, data can be acquired from utility submeters that simply track the energy flowing to one or more chargers. Data collected directly from the electric vehicle supply equipment (EVSE) are typically the most accurate and highest frequency. However, it is not always possible to discern which exact vehicle is being charged during any one session. EVSE-side data can be identified where a single charger ID but a range of vehicle IDs are present (e.g., CH001, EV001-EV005). Data collected from the vehicle’s onboard systems usually does not provide information on which exact charger is being used. Vehicle-side data can be identified where a single Vehicle ID but a range of Charger IDs are present (e.g., EV001, CH001-CH005). Data collected from utility submeters provide no information on which specific vehicle is charging or which specific charger is in use. Submeter data can be identified where multiple Vehicle IDs and multiple Charger IDs are present, but only a single Fleet ID is present (e.g., EV001-EV005, CH001-CH005, Fleet01). The **Charge Data Daily/Session Dictionaries** contains definitions for each available parameter collected as part of an individual charging session, aggregated at either a daily or session level. The parameters available will vary between vehicles and chargers. The **Charger Attributes** table contains specific charger characteristics, coded to at least one anonymous Charger ID and linked to either a single or a range of Vehicle IDs. Vehicle ID can be used as a key between charging data and vehicle attribute tables. The **Charger Attributes Data Dictionary** contains definitions for each available parameter collected on the physical and operational characteristics of the charging hardware itself. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables, and in cases where charging data are supplied, links a vehicle with the charger(s) that supplied it power. The **Charging Data** tables contain the data from each charger’s operations, coded to at least one anonymous Charger ID and linked to either a single or a range of Vehicle IDs. Vehicle ID can be used as a key between charging data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Evaluate data lake design for the accelerator control system

Increasing precision in automation for modern particle accelerators not only creates a requirement to gather data from all devices but also demands scalable and high-performance data infrastructure with the capability of handling vast incoming device data. A well architected data lake is suitable for such a system which integrates real-time data acquisition, transient data caching, and long-term storage. This paper evaluates data lake architecture for an Accelerator Control System (ACS), focusing on two critical components of a data lake, data cache and long-term storage.

Jaikar, Amol [Fermilab]

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES

Using Temporal Information from Human Mobility Data to Detect Anchor Points

Spatiotemporal mobility data are available in massive quantities, but large quantities of data typically include fewer variables or data fields. Often, the only available fields are User ID, Longitude, Latitude, Timestamp (ULLT). This raises an important question: how much can we infer about human mobility patterns using only these four fields? With ULLT data, we do not know individuals' socioeconomic status information or when they are visiting their anchor points (AP) or locations (such as homes, places of employment, or schools), and it is a modern challenge to use this data to infer these characteristics. When detecting anchor locations with limited input information, verification and validation (VV) are significant challenges. This paper addresses the problem of identifying individuals' anchor locations using only temporal information from spatiotemporal datasets with limited attributes. Our approach does not explicitly use latitude and longitude during analysis. Locationbased information is only employed in the preprocessing stage to identify periods of movement (trips) and stops (dwelling). Beyond this step, all analysis is based on temporal patterns. In theory, if stops and dwell times could be detected through alternative means, our method could function entirely without location-based input. We demonstrate this methodology on the 2017 National Household Travel Survey (NHTS) data, because it includes a carefully designed and collected time use survey with representative sampling and labeled ground truth. The high-quality survey data allows us to test the accuracy of our methods because NHTS contains intended place labels and agent/user characteristics. We have also applied our validated AP identification algorithm on very large-scale GPS based trajectory data for Patterns-of-Life (PoL) assessment and other applications, but due to space limit that could not be presented here.

McBride, Liz [ORNL] (ORCID:0000000286925869)