Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Adaptive elasticity policies for staging-based in situ visualization

In situ processing aims to alleviate the growing gap between computation and I/O capabilities by performing data processing close to the data source. In situ processing is widely used to process data generated by multiple data sources, including observation data from edge devices or scientific observational facilities and the simulation data generated by scientific computation on a high-performance computing (HPC) platform. For a scientific workflow that is run on an HPC platform and composed of a simulation program and an in situ data analytics or visualization (abbreviated as ana/vis) task, there is an implicit assumption that the computing resources assigned to the workflow keep static during the workflow execution. However, with the converging trend between the HPC and cloud computing platform, running the in situ ana/vis task in an elastic way is promising to decrease its overhead and improve its resource utilization rate. Resource elasticity represents the ability to change resource configurations such as the number of computing nodes/processes during workflow execution. An elastic job may dynamically adjust resource configurations; it may use a few resources at the beginning and more resources toward the end of the job when interesting data appear. However, it is hard to predict a priori how many computing nodes/processes need to be added/removed during the workflow execution to adapt to changing workflow needs. How to efficiently guide elasticity operations, such as growing or shrinking the number of processes used for in situ analysis during workflow execution, is an open-ended research question. In this article, we present adaptive elasticity policies that adopt workflow runtime information collected during workflow execution to predict how to trigger the addition/removal of processes in order to minimize in situ processing overhead. Taking in situ visualization tasks as an example, we integrate the presented elasticity policies into a staging-based elastic workflow and evaluate its efficiency in multiple elasticity scenarios. Compared with the situation without elasticity or with a static elasticity policy that uses a fixed number of processes for each rescaling operation, the adaptive elasticity policy can save overhead in finding a proper resource configuration and improve resource utilization efficiency. Furthermore, one experiment illustrates that the adaptive elasticity policy saves 41% of core-hours compared with the situation without the resource elasticity.

97 MATHEMATICS AND COMPUTING↗

Open-Source Framework for Data Storage and Visualization of Real-Time Experiments

Digital real time simulators (DRTS) are increasingly being used for the evaluation of power hardware and controller hardware in the laboratory prior to field deployment. Although DRTS are capable of simulating large models in real-time, it is challenging to visualize the results of large models without overdrawing or confounding the viewer. This paper provides an open-source framework for users to visualize their DRTS-based hardware-in-the-loop (HIL) experimental results in real-time. This proposed framework can be used by experimental test beds that can push data through an internet protocol based network. The proposed framework includes three main components. First, it includes the DRTS that generates and pushes the data to a relay. Second, it includes an application that serves multiple purposes, from data storage, testing, and translation of the data to a publisher/subscriber protocol. Finally, it includes libraries and applications that can be used to visualize the data by subscribing to the relay. This framework is available in open source, and it is tested using the HIL platform developed for testing advanced distribution management systems.

advanced distribution management systems↗

A Data-Fusion Method using Bayesian Approach to Enhance Raw Data Accuracy of Position and Distance Measurements for Connected Vehicles

Accurate positioning of vehicles is a critical element of autonomous and connected vehicle systems. Most of other studies heavily focused on enhancing simultaneous localization and mapping (SLAM) methods, i.e., constructing or updating a map of an unknown environment and tracking an object within the map. This paper provides a method that can, in addition to existing SLAM or relevant methods, enhance the raw measurements of position and distance. The basic idea of this study is to identify and update the error distribution of each data source by combining all available information. A Bayesian approach was incorporated to estimate and update the error distribution of individual data sources or sensors. The proposed method can be conducted in real-time environments, and a self-learning scheme determines whether enough data has been collected to further improve the accuracy of such measurements. The simulated experiments show that the proposed model noticeably improves the accuracy of position and distance measurements. Especially, the estimated biases of position coordinates and distance measures are very close to the biases of true error distributions, with the R-squared over 0.98. A similar approach can also be utilized to enhance accuracy of other sensors or measurements in connected vehicle or relevant systems, where multi-data sources are available.

Lim, Hyeonsup↗

Jobs, jobs, jobs: what’s an analyst to do?

Analysts and economists often face the task of using employment metrics to characterize industries of interest. Some key challenges can be understanding where to find employment metrics, the differences in various employment metrics, and when each metric should be used. This article analyzes a variety of publicly available employment data for the United States and compares these data. A detailed description of the intricacies of each data source is provided, which covers factors such as regionality, industry breakout, periodicity, and the types of jobs included. This article provides several case study examples, using the oil and gas extraction, coal mining, and chemical manufacturing sectors to portray challenges data users may face when developing employment estimates that suit their needs. Data users should be aware of a variety of data sources to understand alternative analysis options when data limitations are present and to determine which data source best meets their needs. Instances may occur in which information from one dataset may be used to help impute missing values.

99 GENERAL AND MISCELLANEOUS↗

2025 Large Load Literature Review

This literature review catalogs more than 90 publications focused on large loads, and groups the documents and resources thematically into 12 categories, (listed below). The 2026 Large Load Literature Review and Data Sources summary reports are available here: https://emp.lbl.gov/publications/2026-large-load-literature-review -Load forecasting -Data sources -Reliability and resource adequacy -Large load interconnection -Demand flexibility -Generation -Co-location -Data center location/infrastructure -Large load tariffs -Policy options -Maps and tools -Design and operations

97 MATHEMATICS AND COMPUTING↗

Integrated Hourly Meteorological Database of 20 Meteorological Stations (1981-2022) for Watershed Function SFA Hydrological Modeling

This dataset contains (a) a script “R_met_integrated_for_modeling.R”, and (b) associated input CSV files: 3 CSV files per location to create a 5-variable integrated meteorological dataset file (air temperature, precipitation, wind speed, relative humidity, and solar radiation) for 19 meteorological stations and 1 location within Trail Creek from the modeling team within the East River Community Observatory as part of the Watershed Function Scientific Focus Area (SFA). As meteorological forcings varied across the watershed, a high-frequency database is needed to ensure consistency in the data analysis and modeling. We evaluated several data sources, including gridded meteorological products and field data from meteorological stations. We determined that our modeling efforts required multiple data sources to meet all their needs. As output, this dataset contains (c) a single CSV data file (*_1981-2022.csv) for each location (20 CSV output files total) containing hourly time series data for 1981 to 2022 and (d) five PNG files of time series and density plots for each variable per location (100 PNG files). Detailed location metadata is contained within the Integrated_Met_Database_Locations.csv file for each point location included within this dataset, obtained from Varadharajan et al., 2023 doi:10.15485/1660962. This dataset also includes (e) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and (f) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. Review the (g) ReadMe_Integrated_Met_Database.pdf file for additional details on the script, methods, and structure of the dataset.The script integrates Northwest Alliance for Computational Science and Engineering’s PRISM gridded data product, National Oceanic and Atmospheric Administration’s NCEP-NCAR Reanalysis 1 gridded data product (through the `RCNEP` R package, Kemp et al., doi:10.32614/CRAN.package.RNCEP), and analytical-based calculations. Further, this script downscales the input data into hourly frequency, which is necessary for the modeling efforts.

54 ENVIRONMENTAL SCIENCES↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Making a Water Data System Responsive to Information Needs of Decision Makers

Evidence-based environmental management requires data that are sufficient, accessible, useful and used. A mismatch between data, data systems, and data needs for decision making can result in inefficient and inequitable capital investments, resource allocations, environmental protection, hazard mitigation, and quality of life. In this paper, we examine the relationship between data and decision making in environmental management, with a focus on water management. We focus on the concept of decision-driven data systems —data systems that incorporate an assessment of decision-makers' data needs into their design. The aim of the research was to examine the process of translating data into effective decision making by engaging stakeholders in the development of a water data system. Using California's legislative mandate for state agencies to integrate existing water and other environmental data as a case study, we developed and applied a participatory approach to inform data-system design and identify unmet data needs. Using workshops and focused stakeholder meetings, we developed 20 diverse use cases to assess data sources, availability, characteristics, gaps, and other attributes of data used for representative decisions. Federal and state agencies made up about 90% of the data sources, and could readily adapt to a federated data system, our recommended model for the state. The remaining 10% of more-specialized data, central to important decisions across multiple use cases, would require additional investment or incentives to achieve data consistency, interoperability, and compatibility with a federated system. Based on this assessment, we propose a typology of different types of data limitations and gaps described by stakeholders. We also propose technical, governance, and stakeholder engagement evaluation criteria to guide planning and building environmental data systems. Data-system governance involving both producers and users of data was seen as essential to achieving workable standards, stable funding, convenient data availability, resilience to institutional change, and long-term buy-in by stakeholders. Our work provides a replicable lesson for using decision-maker and stakeholder engagement to shape the design of an environmental data system, and inform a technical design that addresses both user and producer needs.

Cantor, Alida↗

Comparison of greenhouse gas emission estimates from six hydropower reservoirs using modeling versus field surveys

As with most aquatic ecosystems, reservoirs play an important role in the global carbon (C) cycle and emit greenhouse gases (GHG) as carbon dioxide (CO 2 ) and methane (CH 4 ). However, GHG emissions from reservoirs are poorly quantified, especially in temperate systems, resulting in high uncertainty. We compared reservoir C emission estimates and uncertainty of diffusive, ebullitive, and degassing pathways in six hydropower reservoirs in the southeastern United States among four data sources: two field-based surveys and two models (including the GHG Reservoir “G-res” Tool). We found that CH 4 diffusion was most similar across data sources (modeled minus observed, bias = - 21 g CO 2-eq m -2 y -1 ) and had low relative uncertainty (coefficient of variation, CV = 0.98). On the other hand, CO 2 diffusion was least consistent across data sources (bias = - 518 g CO 2-eq m -2 y -1 ). Both field surveys indicated strong negative CO 2 diffusion (i.e., CO 2 uptake) at all reservoirs, while G-res estimated positive CO 2 diffusion. By extension, total C emissions showed similar discrepancies, leading to high uncertainty in upscaling and interpreting reservoir source-sink dynamics. Finally, CH 4 ebullition had the highest relative uncertainty (CV = 2.77) due to high variability across sites. We discuss limitations of field surveys and these models, including temperature-based annualization methods, varying definitions of ebullition zones, low sampling resolution, and lack of dynamism. Future field efforts focused on capturing variability in CO 2 diffusion and CH 4 ebullition will be especially valuable in reducing uncertainty and improving models to advance our understanding reservoir GHG emissions.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Data Catalog Software for Hanford Site Environmental Datasets

Environmental information and data underpin achievement of the U.S. Department of Energy (DOE) Office of Environmental Management (EM) mission at the Hanford Site. The Hanford Environmental Data Management (HEDM) Program is the DOE Richland Operations Office (RL) approach to develop and implement a formal program for managing environmental data and the associated records, materials, and systems at the Hanford Site. The current project, contract, organization, and contractor-specific efforts at managing environmental data sets are insufficient to provide orderly, long-term, site-wide access. A vital element to be created within the HEDM program plan is a catalog of data sources, called the Hanford Environmental Information and Data Index (HEIDI), that will enable long-term access and retrievability for the multiple independent sources of data that might otherwise be difficult to discover. This report compares leading open source and commercial data catalog platforms using criteria to assess the functionality needed to develop the HEIDI catalog of Hanford data sources that connects and exchanges data with established Hanford Local Area Network (HLAN) enterprise information technology systems. Proprietary platforms evaluated included ArcGIS Enterprise Sites, Junar, OpenDataSoft, and Socrata, and non-proprietary platforms included Energy Data eXchange (EDX), Comprehensive Knowledge Archive Network (CKAN), and DKAN (a Drupal-based open data portal based on CKAN). Capabilities supporting data discoverability, retrieval, and archival, as well as metadata standard requirements and integration into the HLAN were rated as either failing to meet requirements (F), meeting requirements (M), or exceeding requirements by delivering additional desired features (E). The lowest rating for any capability area was assigned as the overall rating for the platform. These findings enable DOE-RL and the contractors implementing the HEDM plan to focus on candidate tools likely to meet the requirements for implementing HEIDI. All of the platforms receiving an overall rating of ‘F’ were unable to be deployed on Hanford infrastructure or within dedicated cloud resources. A propriety software-as-a-service (SaaS) model of delivering a data catalog (e.g., found in software such as Junar and OpenDataSoft) favors consistency across customers at the expense of customization and configurable roles that are needed for Hanford work. Hosting data on a shared commercial platform places limits on dataset size (maximum of 240 Mb for OpenDataSoft), a significant limitation for HEIDI implementation. EDX, a government data catalog based on CKAN, received the ‘F’ rating due to an inability to incorporate authentication from HLAN into the system. Among platforms rated ‘M’ or ‘E’, only the Socrata platform had a SaaS delivery model. In contrast to other SaaS platforms, Socrata provided custom roles and gateways that allow local datasets to be incorporated into an online catalog. Socrata also complies with the Federal Risk and Authorization Management Program, a significant benefit for cloud-based management of Hanford data. The other platforms rated ‘M’ or ‘E’, ArcGIS Enterprise Sites, CKAN, and DKAN, provide fully self-hosted options, allowing for greater control and flexibility with the HEIDI catalog. These widely used tools have supportive communities of practice, extensive customization options, and demonstrated deployments that provide evidence that they can meet requirements, often deliver additional desired features, and work well with federal government systems. Completely customized alternatives built on a collection of applications were not evaluated because achieving similar performance to CKAN or DKAN requires substantial resources, especially in the absence of the active communities that have grown to support these tools. ArcGIS Enterprise Sites, Socrata, CKAN, and DKAN were evaluated as strong candidates for successful implementation with HEIDI.

54 ENVIRONMENTAL SCIENCES↗

Development of an open-source regional data assimilation system in PEcAn v. 1.7.2: application to carbon cycle reanalysis across the contiguous US using SIPNET

Abstract. The ability to monitor, understand, and predict the dynamics of the terrestrial carbon cycle requires the capacity to robustly and coherently synthesize multiple streams of information that each provide partial information about different pools and fluxes. In this study, we introduce a new terrestrial carbon cycle data assimilation system, built on the PEcAn model–data eco-informatics system, and its application for the development of a proof-of-concept carbon “reanalysis” product that harmonizes carbon pools (leaf, wood, soil) and fluxes (GPP, Ra, Rh, NEE) across the contiguous United States from 1986–2019. We first calibrated this system against plant trait and flux tower net ecosystem exchange (NEE) using a novel emulated hierarchical Bayesian approach. Next, we extended the Tobit–Wishart ensemble filter (TWEnF) state data assimilation (SDA) framework, a generalization of the common ensemble Kalman filter which accounts for censored data and provides a fully Bayesian estimate of model process error, to a regional-scale system with a calibrated localization. Combined with additional workflows for propagating parameter, initial condition, and driver uncertainty, this represents the most complete and robust uncertainty accounting available for terrestrial carbon models. Our initial reanalysis was run on an irregular grid of ∼ 500 points selected using a stratified sampling method to efficiently capture environmental heterogeneity. Remotely sensed observations of aboveground biomass (Landsat LandTrendr) and leaf area index (LAI) (MODIS MOD15) were sequentially assimilated into the SIPNET model. Reanalysis soil carbon, which was indirectly constrained based on modeled covariances, showed general agreement with SoilGrids, an independent soil carbon data product. Reanalysis NEE, which was constrained based on posterior ensemble weights, also showed good agreement with eddy flux tower NEE and reduced root mean square error (RMSE) compared to the calibrated forecast. Ultimately, PEcAn's new open-source regional data assimilation framework provides a scalable workflow for harmonizing multiple data constraints and providing a uniform synthetic platform for carbon monitoring, reporting, and verification (MRV) as well as accelerating terrestrial carbon cycle research.

54 ENVIRONMENTAL SCIENCES↗

WA-Omic_LA.1.0 - Quantitative Lipidomics, Metabolomics, and Sequencing (16S/ITS) Publication Data DOI Package

Corresponding Data Publication: "Rapid remodeling of the soil lipidome in response to a drying-rewetting event." This study reveals specific changes in lipids and metabolites that are indicative of stress adaptation, substrate use, and cellular recovery during soil drying and subsequent rewetting. Drought induced nutrient limitation was reflected in the lipidome and polar metabalome, both of which rapidly shifted (within hours) upon rewet. Reduced nutrient access in dry soil caused the replacement of glycerophospholipids with phosphorus-free lipids and impeded resource-expensive osmolyte accumulation. Elevated levels of ceramides and lipids with long chain polyunsaturated fatty acids, in dry soil suggests that lipids play an important role in fungal drought tolerance. Increasing abundance of bacterial glycerophospholipids and triacylglycerols with fatty acids typical of bacteria and polar metabolites suggest metabolic recovery in representative bacteria once the environmental conditions are conducive for growth. These results underscore the importance of the soil lipidome as a robust indicator of microbial community responses, especially at the short time scales of cell-environment reactions. Data package contents reported here are the first version and contain pre- and post-processed data acquisition and subsequent downstream analysis files using various data source instrument method techniques and Mass Spectroscopy (MS) EMSL capabilities. This publication data package DOI is a comprehensive high-throughput multi-omics data lifecycle collection containing processed data method metadata. Support files include additional data download “Read Me” file containing data descriptor information and data source application ontologies (see data dictionary). Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. For increased data availability and interoperability, GC-MS/LC-MS mass spectrometry datasets (Thermo .raw ) were deposited at the MassIVE database repository under the related data accession MSV000086931 and can be accessed by using the API. Statistical data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location.

Amplicon sequencing 16S ITS LC-MS/MS lipidomics mu↗

WA-Omic_LA.1.0 - Quantitative Lipidomics, Metabolomics, and Sequencing (16S/ITS) Publication Data DOI Package

Corresponding Data Publication: "Rapid remodeling of the soil lipidome in response to a drying-rewetting event." This study reveals specific changes in lipids and metabolites that are indicative of stress adaptation, substrate use, and cellular recovery during soil drying and subsequent rewetting. Drought induced nutrient limitation was reflected in the lipidome and polar metabalome, both of which rapidly shifted (within hours) upon rewet. Reduced nutrient access in dry soil caused the replacement of glycerophospholipids with phosphorus-free lipids and impeded resource-expensive osmolyte accumulation. Elevated levels of ceramides and lipids with long chain polyunsaturated fatty acids, in dry soil suggests that lipids play an important role in fungal drought tolerance. Increasing abundance of bacterial glycerophospholipids and triacylglycerols with fatty acids typical of bacteria and polar metabolites suggest metabolic recovery in representative bacteria once the environmental conditions are conducive for growth. These results underscore the importance of the soil lipidome as a robust indicator of microbial community responses, especially at the short time scales of cell-environment reactions. Data package contents reported here are the first version and contain pre- and post-processed data acquisition and subsequent downstream analysis files using various data source instrument method techniques and Mass Spectroscopy (MS) EMSL capabilities. This publication data package DOI is a comprehensive high-throughput multi-omics data lifecycle collection containing processed data method metadata. Support files include additional data download “Read Me” file containing data descriptor information and data source application ontologies (see data dictionary). Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. For increased data availability and interoperability, GC-MS/LC-MS mass spectrometry datasets (Thermo .raw ) were deposited at the MassIVE database repository under the related data accession MSV000086931 and can be accessed by using the API. Statistical data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location.

Amplicon sequencing 16S ITS LC-MS/MS lipidomics mu↗

WA-Omic_LA.1.0 - Quantitative Lipidomics, Metabolomics, and Sequencing (16S/ITS) Publication Data DOI Package

Corresponding Data Publication: "Rapid remodeling of the soil lipidome in response to a drying-rewetting event." This study reveals specific changes in lipids and metabolites that are indicative of stress adaptation, substrate use, and cellular recovery during soil drying and subsequent rewetting. Drought induced nutrient limitation was reflected in the lipidome and polar metabalome, both of which rapidly shifted (within hours) upon rewet. Reduced nutrient access in dry soil caused the replacement of glycerophospholipids with phosphorus-free lipids and impeded resource-expensive osmolyte accumulation. Elevated levels of ceramides and lipids with long chain polyunsaturated fatty acids, in dry soil suggests that lipids play an important role in fungal drought tolerance. Increasing abundance of bacterial glycerophospholipids and triacylglycerols with fatty acids typical of bacteria and polar metabolites suggest metabolic recovery in representative bacteria once the environmental conditions are conducive for growth. These results underscore the importance of the soil lipidome as a robust indicator of microbial community responses, especially at the short time scales of cell-environment reactions. Data package contents reported here are the first version and contain pre- and post-processed data acquisition and subsequent downstream analysis files using various data source instrument method techniques and Mass Spectroscopy (MS) EMSL capabilities. This publication data package DOI is a comprehensive high-throughput multi-omics data lifecycle collection containing processed data method metadata. Support files include additional data download “Read Me” file containing data descriptor information and data source application ontologies (see data dictionary). Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. For increased data availability and interoperability, GC-MS/LC-MS mass spectrometry datasets (Thermo .raw ) were deposited at the MassIVE database repository under the related data accession MSV000086931 and can be accessed by using the API. Statistical data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location.

Amplicon sequencing 16S ITS LC-MS/MS lipidomics mu↗

Energy Emergency and Preparedness Data: Frequently Asked Questions (FAQs) and Quick Guidance on Crisis Communications

State Energy Offices and Public Utility Commissions rely on timely, accurate, and actionable information to perform their energy emergency response duties and execute their roles as state energy security planners. In support of this need, the National Association of State Energy Officials (NASEO) and the National Association of Regulatory Utility Commissions (NARUC) hosted an Energy Security and Data Analysis Workshop in Washington, DC to identify energy security response and planning data sources; and to share successful methods of data use and integration in state, federal, and private sector tools. Following the workshop, NASEO and NARUC hosted two topical data-centric webinars based on state priorities to identify best practices in Geographic Information Systems (GIS) and Crisis Communications programs leveraged in energy assurance planning and response. The Crisis Communications webinar covered best practice tactics for how states can respond during energy emergencies and other crises. Based on the workshop and webinars, this document summarizes commonly used data sources and includes frequently asked questions which may be used by state energy officials (i.e., consisting of staff from Public Utility Commissions and Governor-designated State Energy Offices) to help guide them in developing or improving their Crisis Communications capabilities. A strong public information program is a key crisis management tool. Timely and accurate information helps prevent confusion and uncertainty and encourages public support and cooperation. As energy subject matter experts, state energy officials are vital in the distilling, clarifying, and conveying energy sector information and implications to decision-makers and the public. Other participants in an effective public information program include the Governor’s Office, other state agencies, local governments, energy providers, local businesses, state legislature, and the federal government. It is essential to provide stakeholders and the public with information about the nature, severity, and duration of an emergency because inadequate understanding and awareness can lead to undesirable actions that could further exacerbate the situation. Before a state government can provide information to the public, it must gather information, describe the emergency accurately, and develop recommendations to manage the situation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Modeling freight mode choice using machine learning classifiers: a comparative study using Commodity Flow Survey (CFS) data

This study explores the usefulness of machine learning classifiers for modeling freight mode choice. We investigate eight commonly used machine learning classifiers, namely Naïve Bayes, Support Vector Machine, Artificial Neural Network, K-Nearest Neighbors, Classification and Regression Tree, Random Forest, Boosting and Bagging, along with the classical Multinomial Logit model. US 2012 Commodity Flow Survey data are used as the primary data source; we augment it with spatial attributes from secondary data sources. The performance of the classifiers is compared based on prediction accuracy results. The current research also examines the role of sample size and training-testing data split ratios on the predictive ability of the various approaches. In addition, the importance of variables is estimated to determine how the variables influence freight mode choice. The results show that the tree-based ensemble classifiers perform the best. Specifically, Random Forest produces the most accurate predictions, closely followed by Boosting and Bagging. With regard to variable importance, shipment characteristics, such as shipment distance, industry classification of the shipper and shipment size, are the most significant factors for freight mode choice decisions.

42 ENGINEERING↗

Multi-Kernel Adaptive Support Vector Machine for Scalable Predictive Maintenance

Application of data-driven solutions across an industry is challenging, since the data are often stored locally, and increasing privacy and security concerns restrict access to the data. In addition, it is highly unlikely that all potential data patterns are captured in a single data source. Because it is highly unlikely that all potential data patterns are captured in a single data source, machine learning (ML) models developed from a single source cannot be robust enough. An alternative is to train the ML model at each source and develop a distributed knowledge discovery and aggregation approach to build global knowledge. In this paper, we develop and demonstrate a distributed ML model, federated transfer learning (FTL), using a multi-kernel-based adaptive support vector machine (MK-A-SVM). For federated learning (FL), the multi-kernel (MK) approach enables feature-specific model aggregation under data heterogeneity; whereas for transfer learning (TL) the adaptive model enables utilization of an aggregated model from a different task. The proposed approach is validated using nuclear power plant (NPP) vertical motor-driven pump data to predict the health condition of vertical motor-driven pumps as an anomaly detection. The efficiency of the proposed approach is also quantified and compared with neural network.

42 ENGINEERING↗

Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins

The Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins contains a series of spatial datasets representing spatial extents of publicly available data for caprock and seal rock units within the Appalachian Basin, Denver-Julesburg Basin, Great Valley Basin (Sacramento and San Joaquin Basins), Illinois Basin, Michigan Basin, San Juan Basin, U.S. Gulf Coast Basin, and Williston Basin. The database is designed to support carbon storage feasibility and resources assessment for carbon transport and storage (CTS) projects while displaying the spatial extent of prospective seal units and provide a guide to the original data source. This database leverages publicly available data resources from authoritative sources (e.g. U.S. Geological Survey, State Geologic Surveys, and published reports), and aims to help guide users to understand the seal unit's spatial coverage and data gaps from the regional to sub-basin/field scale. The database is organized by seal unit/formation, including the spatial extent for data found to be available for the seal unit. The various datasets represented include spatial extents of the lithologic formation, depth to top structural contour maps, and thickness/isopach maps. Included in this submission are the following resources: 1. Geodatabase/Dataset: “prospective-seal-unit-extents-2025.gdb” 2. ReadMe: “readme-prospective-seal-unit-spatial-extent-dataset-2025.pdf” 3. Data Catalog: “prospective-seal-unit-spatial-extents-data-catalog-2025.xlsx” 4. Data Sources Key: “data-source.csv” Please see NETL disclaimers here: https://netl.doe.gov/home/disclaimer

Basin↗