Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data platform”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

520EV Refuse Truck Telemetry Dataset

This dataset contains telemetry data for the 520EV refuse trucks accessed through a Veracity data platform as provided by Peterbilt. The dataset contains driving data on both electric trucks used by SWS. Data were recorded at an hourly resolution and contain energy use data while driving and idling, distance driven (miles), and driving speed (miles per hour). This dataset also contains charging data on the fast charger in kilowatt-hours for each hour.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)

Accelerated data-driven materials science with the Materials Project

The Materials Project was launched formally in 2011 to drive materials discovery forwards through high-throughput computation and open data. More than a decade later, the Materials Project has become an indispensable tool used by more than 600,000 materials researchers around the world. This Perspective describes how the Materials Project, as a data platform and a software ecosystem, has helped to shape research in data-driven materials science. We cover how sustainable software and computational methods have accelerated materials design while becoming more open source and collaborative in nature. Next, we present cases where the Materials Project was used to understand and discover functional materials. We then describe our efforts to meet the needs of an expanding user base, through technical infrastructure updates ranging from data architecture and cloud resources to interactive web applications. Finally, we discuss opportunities to better aid the research community, with the vision that more accessible and easy-to-understand materials data will result in democratized materials knowledge and an increasingly collaborative community.

Horton, Matthew K

Data for A Generalized Platform for Artificial Intelligence-powered Autonomous Protein Engineering

Proteins are the molecular machines of life with numerous applications in energy, health, and sustainability. However, engineering proteins with desired functions for practical applications remains slow, expensive, and specialist-dependent. Here we report a generally applicable platform for autonomous enzyme engineering that integrates machine learning and large language models with biofoundry automation to eliminate the need for human intervention, judgement, and domain expertise. Requiring only an input protein sequence and a quantifiable way to measure fitness, this automated platform can be applied to engineer a wide array of proteins. As a proof of concept, we engineer Arabidopsis thaliana halide methyltransferase (AtHMT) for a 90-foldimprovement in substrate preference and 16-fold improvement in ethyl-transferase activity, along with developing a Yersinia mollaretii phytase (YmPhytase) variant with 26-fold improvement in activity at neutral pH. This is accomplished in four rounds over 4 weeks, while requiring construction and characterization of fewer than 500 variants for each enzyme. This platform for autonomous experimentation paves the way for rapid advancements across diverse industries, from medicine and biotechnology to renewable energy and sustainable chemistry.

AI/ML

Equitable Energy Metrics for Integration into Building Performance Standard Tracking Platforms: Preprint

Building Performance Standards (BPS) are being adopted globally and in the United States of America, where 14 different states and jurisdictions have a policy in place and many others are under development (Department of Energy (DOE) 2023). Accurate and equitable data sources are essential to make informed decisions about focusing investment on upgrading buildings to meet jurisdictional goals. There have been multiple new tools developed related to Energy Equity and Environmental Justice (EEEJ) and the resulting datasets need to be integrated into large building port-folios for quick access and better scalability. Integrating EEEJ data in a user-friendly format can help decision makers more quickly assess impacts and analyze the multitude of potentially significant metrics for which there is not yet consensus. In the U.S. and Canada, many BPS ordinances rely primarily on ENERGY STAR Portfolio Manager (ESPM) to capture building characteristics and energy and water consumption data. These datasets can then be imported into city-specific building tracking tools like the Standard Energy Efficiency Data Platform (SEED). Crucially, BPS decision makers require an efficient means of identifying buildings in priority communities to effectively allocate resources and funding. This process must integrate seamlessly with existing jurisdictional toolsets for optimal utility. This paper will demonstrate, for the case of Washington D.C.'s (the District) data, a workflow that provides actionable data for building upgrade investment prioritization in disadvantaged communities.

BPS

Data mining the missing ordered phases of Li/Na metal oxides

Data-driven discovery of Li-ion and Na-ion battery materials has been pioneered by generic materials data platforms such as the Materials Project. After decades of progress, it is timely to ask whether there remain underexplored compositional spaces. Here, in this work, we present a systematic data-mining effort to uncover missing ordered binary, ternary and quaternary Li/Na-containing metal oxides using high-throughput density functional theory (DFT). Building on 19,120 stable and metastable oxides entries from the Materials Project, we performed 13,245 additional calculations through isovalent substitutions of known ground states, experimentally reported compounds, and specific prototype structures. Our study identifies 36 new ground states within the GGA/GGA + U convex hull and 45 within the r 2 SCAN convex hull. Additionally, we identified 840 metastable compounds from GGA/GGA + U and 979 from r 2 SCAN that are absent in the present Materials Project databases. Moreover, we have tripled the metastable materials in compositional spaces with a molar ratio of cation/anion >1, highlighting the overlooked opportunities in this compositional space.

25 ENERGY STORAGE

A path to intelligent watersheds: coordinating the data to decision pipeline

Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.

Colorado River

Metallic fuel transient fuel-cladding interface liquefaction model assessment platform enabled by integrating BISON with databases

A novel platform has been developed within the BISON fuel performance code to assess models of fuel-cladding interface liquefaction for sodium-cooled fast reactor (SFR) metallic fuels. Here, this platform is crucial because liquefaction at the fuel-cladding interface significantly impacts fuel performance and may compromise fuel pin integrity during transient events. To ensure accurate predictions, the platform integrates data collected during the Integral Fast Reactor (IFR) program, now archived in metallic fuel databases. This integration supports verification and validation (V&V) of the models in BISON. Leveraging the extensive US experience with metallic fuel liquefaction and the collections of preserved legacy data, the platform serves as a powerful tool for evaluating existing models and advancing the development of new ones.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Preliminary Evaluation of Joint Electricity-Hydrogen Concept of Operations

An initial thermal power dispatch (TPD) concept of operations was evaluated that couples a nuclear power plant to a nearby hydrogen production plant. GSE Systems’ generic pressurized water reactor full-scope simulator was modified with a TPD model comprised of a thermal power extraction and delivery system. A prototype human-system interface (HSI) was developed to interact with the TPD model and allow participants to execute the basic operating scenarios for normal operations. Four retired operators performed the evaluation, and due to COVID-19 travel restrictions, the original in-person experimental design was restructured to support a remote participator evaluation using a web meeting platform. Data from operator feedback, observations from the research team, and quantitative survey responses revealed that the initial TPD concept of operations is feasible. The operators were comfortable with the engineered system and HSI and could manage it without adverse impacts to reactor power, plant safety, or equipment. Findings are discussed in terms of both the TPD system design and HSI performance.

human factors

Postearthquake Damage Mapping via Remote Sensing: Lessons From the 2023 Türkiye Disaster

This review addresses the urgent need for scalable, accurate, and reproducible remote sensing solutions following the February 2023 Türkiye earthquakes. It synthesizes the contributions of five peer-reviewed studies published in the IEEE JSTARS Special Issue on postearthquake damage and risk assessment. These studies cover areas such as damage classification with deep learning, fusion of multisource remote sensing data, creation of benchmark datasets, detailed damage mapping, and analysis of geophysical signals using outgoing longwave radiation. The article summarizes the methodological approaches and the practical relevance of the reviewed studies for detecting, evaluating, and quantifying damage, and outlines key challenges, including model generalization, class ambiguity, and data integration. It also discusses emerging trends, including explainable artificial intelligence, multimodal data fusion, and open-data platforms. This synthesis provides a foundation for building robust, interpretable, and real-time disaster response systems and aims to guide future research in earthquake-related Earth observation and rapid damage assessment.

Taskin, Gulsen [Istanbul Technical University] (OR

FleetREDI Insight: Beverage Delivery in New York City

Capturing real-world data is critical to improving efficiency and supporting technology advancements in commercial vehicles. FleetREDI’s insights provide detailed duty cycle information and highlight unique aspects of the given dataset. Each insight delivers a quick look at the collected data by summarizing the operation and identifying key findings of the initial analysis. This FleetREDI insight explores beverage delivery tractors operating in New York City. Last-mile beverage delivery supports local bars and restaurants throughout Manhattan and the broader New York City area. Manhattan Beer Distributors is a beverage delivery company operating in Manhattan and the Bronx. Logging devices were installed in 17 vehicles, and operational data were collected between August and October 2022. Two types of vehicles were included in data collection: 7 tractors and 10 bay trucks. Using NLR’s FleetREDI data platform, this dataset provides a summary of daily operation to help understand duty cycle characteristics. This includes daily distance, fuel use, and estimated engine-produced energy consumption for 17 bay trucks and tractors that operated more than 7,500 miles in slow-speed urban operation. ![FleetREDI beverage delivery](FleetREDI-beverage-delivery-nyc.jpg)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

FleetREDI Insight: Intrastate Coach Bus Dataset

Capturing real-world data is critical to improving efficiency and supporting technology advancements in commercial vehicles. FleetREDI’s insights provide detailed duty cycle information and highlight unique aspects of the given dataset. Each insight delivers a quick look at the collected data by summarizing the operation and identifying key findings of the initial analysis. This FleetREDI insight explores coach buses operating in Colorado. Coach buses are a primary mover for intrastate transit and are primarily used for longer trips with more comfortable seats and a restroom. All Aboard America! Holdings Inc. offers various fixed-service and charter routes across Colorado on its Bustang fleet out of its depot in Golden, Colorado. NLR installed logging devices and collected operational data on nine 40-foot Bustang motorcoaches operating on fixed routes from May through August 2022. Using NLR’s FleetREDI data platform, this dataset provides a summary of daily operation to help understand duty cycle characteristics. This includes daily distance, fuel use, and estimated engine-produced energy consumption for nine motorcoaches that operated more than 33,000 miles. These vehicles primarily operated on Interstate 25 and Interstate 70. ![FleetREDI interstate bus](FleetREDI-interstate-bus.jpg)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS

Adapt: A Weather Radar Data Analysis and Nowcasting Platform for Informed Adaptive Scanning

SF-26-021 Adapt is a data processing platform for real-time data analysis, short term prediction of targets convective cells and tracking for archived data. It provides tools for downloading, processing, segmenting, projecting, analyzing, and visualizing storm cell data from weather radar. The pipeline includes cell detection, motion estimation using optical flow, cell property extraction, and persistence to NetCDF and SQLite/Parquet for guiding adaptive scanning.

Raut, Bhupendra Ashokrao [Argonne National Laborat

Ocelot: An Interactive, Efficient Distributed Compression-As-a-Service Platform With Optimized Data Compression Techniques

Large volumes of data generated by scientific simulations, genome sequencing, and other applications need to be moved among clusters for data collection/analysis. Data compression techniques have effectively reduced data storage and transfer costs. However, users' requirements on interactively controlling both data quality and compression ratios are non-trivial to fulfill. Here, we propose a novel Compression-as-a-Service (CaaS) platform called Ocelot with four important contributions: (1) It offers real-time visualization, interactive compression, and transfer of scientific datasets. (2) It incorporates new strategies for compressing diverse types of datasets more effectively than traditional methods. (3) It provides an effective method for estimating the compression ratio and execution time of compression tasks. (4) Experiments on multiple real-world datasets on geographically distributed computers show that Ocelot can significantly improve data transfer efficiency with a performance gain of more than 10x in computing clusters with relatively slow networks.

compression as a service (CaaS)

A high-throughput experimentation platform for data-driven discovery in electrochemistry

Automating electrochemical analyses combined with artificial intelligence is poised to accelerate discoveries in renewable energy sciences and technologies. This study presents an automated high-throughput electrochemical characterization (AHTech) platform as a cost-effective and versatile tool for rapidly assessing liquid analytes. The Python-controlled platform combines a liquid handling robot, potentiostat, and customizable microelectrode bundles for diverse, reproducible electrochemical measurements in microtiter plates, minimizing chemical consumption and manual effort. To showcase the capability of AHTech, we screened a library of 180 small molecules as electrolyte additives for aqueous zinc metal batteries, generating data for training machine learning models to predict Coulombic efficiencies. Key molecular features governing additive performance were elucidated using Shapley Additive exPlanations and Spearman’s correlation, pinpointing high-performance candidates like cis-4-hydroxy-d-proline, which achieved an average Coulombic efficiency of 99.52% over 200 cycles. The workflow established herein is highly adaptable, offering a powerful framework for accelerating the exploration and optimization of extensive chemical spaces across diverse energy storage and conversion fields.

Lin, Dian-Zhao [Johns Hopkins University, Baltimor

Vulcan Test Platform: Demonstrating the Data Center as a Flexible Grid Asset

Explosive data center demand is outpacing grid infrastructure development. AI workloads and hyperscale cloud growth are creating unprecedented power requirements, while traditional grid expansion faces multi-year development timelines, regulatory hurdles, and decarbonization challenges. Sidewalk Infrastructure Partners recognized this impending crisis years ago and founded Verrus to develop an innovative solution: data centers that function as grid assets rather than passive loads. The Verrus approach integrates proprietary grid-aware controls with battery energy storage systems (BESS) in a medium-voltage architecture that delivers three critical capabilities: Fast-responding demand flexibility that can service requests from the utility within 10 seconds, Uninterrupted transition to islanded operation during grid disturbances, Continuous uptime assurance while maintaining all customer service level commitments Through Verrus' strategic partnership with the National Renewable Energy Laboratory (NREL), we validated these capabilities on Vulcan, a 70 MW utility-scale test platform powered by NREL's ARIES Virtual Emulation Environment. This deployment-ready technology has successfully demonstrated that Verrus data centers can deliver meaningful grid services while maintaining mission-critical reliability. This technical report outlines the design, methodology, and results of this emulated deployment, demonstrating that data centers can play a pivotal role in enhancing grid flexibility and reliability, without sacrificing service level guarantees.

24 POWER TRANSMISSION AND DISTRIBUTION