Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data publication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ComStock Measure Documentation: Variable-Speed Pumps

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - variable speed pumps - and briefly introduces key results. The full public data set can be accessed on the ComStock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ComStock Measure Documentation: High-Efficiency Rooftop Unit

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - high-efficiency rooftop unit (RTU) - and briefly introduces key results. The full public data set can be accessed on the Comstock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Downloadable Dynamometer Database (D3): Public Test Data on Advanced-Technology Vehicles

Access to high-quality, independent vehicle test data is critical to advancing energy-efficient transportation research. The Downloadable Dynamometer Database (D3) is a public repository of dynamometer test data on advanced-technology vehicles, generated at the Advanced Mobility Technology Laboratory (AMTL) at Argonne National Laboratory and hosted by the Transportation and Power Systems Division. The database has been made available to support researchers, students, and professionals engaged in energy-efficient vehicle research, development, and education. A wide range of vehicle categories has been tested (i.e., alternative fuel vehicles, conventional gasoline and diesel vehicles, all-electric vehicles, hybrid electric vehicles, and plug-in hybrid electric vehicles), as well as various drive cycles and test conditions documented in the accompanying D3 user presentation. Stakeholders can select a vehicle type, identify a vehicle of interest, and download the associated test data for use in their own analyses. Data downloaded from D3 must be accompanied by the required attribution: "This data is from the Downloadable Dynamometer Database and was generated at the Advanced Mobility Technology Laboratory (AMTL) at Argonne National Laboratory." These data are critical to vehicle modeling, validation, technology assessment, and educational use.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Quantifying Annual Industrial Locomotive Energy Consumption in the United States

While US Class 1 railroad locomotive rosters and annual fuel consumption are well-documented, considerably less is known regarding the overall energy consumption of operations involving industrial locomotives. To determine the energy savings potential of this rail operating sector, the objective of this research is to develop an inventory of US industrial locomotives and a baseline estimate of their annual energy consumption. Creating an industrial locomotives roster from public data is challenging given their diverse ownership by shippers or leasing companies, and operating locales largely out of public view. By cross-referencing public data on locomotive reporting marks, serial numbers, online images and aerial images, the project team confirmed the age, model and horsepower of over one thousand industrial locomotives. Estimating energy consumption is complicated by the variability in industrial locomotive types and power ratings, and extreme differences in duty cycles and utilization. Given these limitations, using quantified case study examples and adjustments to standard EPA line-haul and switching duty cycles, bounds on the magnitude of annual US industrial locomotive energy consumption were estimated.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Marine Hydrokinetic Tidal Turbine

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset is part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with other energy technologies. This dataset contains inputs and outputs from simulations of a floating marine hydrokinetic turbine over approximately half a tidal cycle (~6.6 hours). Inflow conditions were derived from field measurements in Alaska’s Cook Inlet and represent a tidal environment in which the current speed ramps from near 0 m/s to a peak of 3 m/s and back. The original acoustic doppler current profiler dataset is publicly available on the Marine and Hydrokinetic Data Repository. In a full tidal cycle, the flow reverses and the rotor would reorient; this reversal was not modeled. In the Cook Inlet campaign , turbulence intensity was similar in both directions. Two inflow cases are included. In the first case, labeled “raw” in the files, the measured current time series was used directly in the InflowWind module of OpenFAST. Speed and direction were applied as a function of time and elevation, uniformly in the horizontal direction. With full spatial coherence, this approach captures high turbulent variability and results in pronounced power fluctuations, so it is considered a conservative, near-worst-case representation of loading. In the second case, labeled “average” in the files, a 30-minute moving average was applied to extract the slowly varying mean speed. The residual fluctuations about this mean were used to generate spatially varying, full-field turbulence inputs with TurbSim, giving a more physically realistic representation of the inflow across the rotor disk. Two random realizations were used to produce distinct inflow conditions for two OpenFAST simulations representing a two-turbine array. The same turbulence intensity is applied across the full time series, producing larger fluctuations at the start and end, where the mean speed is low. The second case is the more appropriate framework for performance and power assessment but overpredicts turbulence at lower flow speeds and underpredicts it at higher speeds. As the floating platform moves and the rotor changes its x-position, Taylor’s frozen turbulence hypothesis used by InflowWind assumes a constant rather than a time-varying mean velocity, introducing some inaccuracy in the velocity plane sampling. The turbine modeled is the 500-kW Reference Model 1, a horizontal-axis two-bladed hydrokinetic turbine on a four-column floating semisubmersible substructure . Simulations were performed using OpenFAST v4.1 with the Reference Open Source Controller (ROSCO) v2.10. All input files required to reproduce the simulations are included. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel . This unit supports up to 2.5 MW, but NLR has only a single 1.25-MW stack. The datasets report hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. The system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operating current of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The simulated tidal turbine time series data was translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz. Each zip file represents a single tidal electrolysis experiment and is named: {technology}_{inflow method}_{number of 500 kW tidal turbines connected} For instance, “tidal-500kW-RM1_average_2.zip” is a 6-hour experiment using the 500-kW tidal reference model, scaled by 2x (1-MW) to better match the electrolyzer maximum of 1.25MW, fed with the 30-minute moving average current case. Each zip folder contains the following files: A .csv file of raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. A .csv file combines all tidal profiles as "combined_tidal_experiments.csv." A separate experiment, “characterization_200.zip,” shows the MC250 electrolyzer steady-state response with 30-minute load steps over 5 hours and is accessible with this entry.

08 HYDROGEN↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis – Simulated Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wind turbine. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen . While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the simulated wind energy profiles, NLR used OpenFAST to simulate a 3.4-MW International Energy Agency (IEA) reference wind turbine. The hour-long wind energy profiles varied over wind turbulence intensity (Class A or Class C) and average wind speed (5, 7, or 9 m/s). To match the power limits of the 1.25-MW electrolyzer and 3.4-MW IEA wind turbine most effectively and to maximize the efficiency of hydrogen production at a given average wind speed, the profiles were sometimes scaled by two times. This means that, in some cases, the experimental setup assumed two 1.25-MW electrolyzers were coupled with the wind turbine, representing a total maximum electrolysis load of 2.5 MW. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}-{average wind speed}-{turbulence class}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “windIEA3.4-5ms-C_2-400.zip” represents the hour-long experiment using the IEA 3.4-MW turbine, subjected to an average wind speed of 5 m/s and Class C wind turbulence, and connected to two 1.25-MW electrolyzers with the power supply set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wind turbine power. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis .

08 HYDROGEN↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Wave

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production using a single, simulated wave energy conversion device. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel Hydrogen. While the unit supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. For the wave energy, NLR used a wave energy converter model from PacWave. These devices can be equipped with accumulators and pressure relief values to smooth the power output by storing and releasing hydraulic energy. Using a peak power output of 10 MW, the model created two 25-minute profiles: one with and one without the accumulators and pressure relief valves. To down select the profile data from the native resolution of 20 Hz to 1 Hz, NLR took the mean of every 20 data points. NLR experimented with two simulated wave energy power plants: one that peaks at 10 MW, and one that peaks at 5 MW. These profiles were scaled for the physical 1.25 MW electrolyzer by multiplying the original profiles by one eighth and one quarter, respectively. The first profile matches the capacity rating of eight of the 1.25 MW electrolyzers, while the second matches four electrolyzers. Finally, NLR experimented with two settings for the electrolyzer power supply minimum and maximum current ramp rates (gain and slew): 200 and 400 amperes per second. The simulated profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wave electrolysis experiment and is formatted as follows: {technology}-{accumulator?}_{number of 1.25 MW electrolyzers connected}-{electrolyzer ramp rate in amperes/second} For instance, “wavePacWave-Noacc_4-400.zip” represents the 25 minute-long experiment using the PacWave’s wave energy converter model, equipped with no accumulator, connected to four 1.25-MW electrolyzers with their power supplies set to a maximum current ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. An experiment, labeled “characterization_200.zip”, demonstrates the MC250 electrolyzer steady-state response with 30 minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all wave profiles combined into one dataset labeled "combined_wave_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis.

08 HYDROGEN↗

NNFDivergence

The code implements f divergence regularization for neural networks in the Python-based Pytorch framework. The methods are the main focus but the repository will also contain examples that operate on purely synthetic "toy" data or on openly available, public data from NASA.

Klein, Natalie [@lanl]↗

Electricity Baseline 2022 Background Data and Log File

The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2022 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilized the appdirs Python dependency (https://pypi.org/project/appdirs/). This submission includes the background data used to generate the 2022 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: `python -c "import appdirs; print(appdirs.user_data_dir())"`). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2022 model run is also included, which contains the statements at the DEBUG level and above.

Electricity; LCA; data inventory↗

Electricity Baseline 2021 Background Data and Log File

The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2021 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilizes the appdirs Python dependency (https://pypi.org/project/appdirs/). An overview of the ElectricityLCI data stores may be found on the README (https://github.com/USEPA/ElectricityLCI/blob/v2.0/README.md#data-store). This submission includes the background data used to generate the 2021 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: python -c "import appdirs; print(appdirs.user_data_dir())"). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2021 model run is also included, which contains the statements at the DEBUG level and above.

Electricity; LCA; LCI; Life Cycle; data inventory↗

Electricity Baseline 2020 Background Data and Log File

The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2020 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilizes the appdirs Python dependency (https://pypi.org/project/appdirs/). An overview of the ElectricityLCI data stores may be found on the README (https://github.com/USEPA/ElectricityLCI/blob/v2.0/README.md#data-store). This submission includes the background data used to generate the 2020 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: python -c "import appdirs; print(appdirs.user_data_dir())"). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2020 model run is also included, which contains the statements at the DEBUG level and above.

Electricity; LCA; LCI; Life Cycle; data inventory↗

Unlocking nighttime mobility: Land use and accessibility in public transit for night commuters

Night commuters are integral to urban transportation systems. Essential services such as healthcare and manufacturing rely on workers who travel at night, and reliable mobility options are crucial for them. A gap exists in understanding how land use and accessibility influence public transportation use among night commuters. This study addresses this gap by using public data to explore land use and accessibility factors that affect night commuters' public transportation use in New York State. We investigated (1) the demographic characteristics of night commuters; (2) the influence of land use and accessibility on nighttime public transportation use; and (3) potential improvements to increase public transportation use and their impact. We combined data from the National Household Travel Survey with the Smart Location Database to link home locations with land use characteristics. Using logistic regression, we found that although females are generally less likely to be night commuters, they are more likely to use public transportation. Longer commute distances are associated with higher use of public transportation. Increasing job density along fixed-guideway transit routes and improving overall job accessibility via public transportation significantly enhances public transportation use among night commuters. In conclusion, this research provides actionable insights for public transportation agencies and urban planners to support night commuters, improving access and encouraging nighttime employment.

Job accessibility↗

Sharing the Sun: Community Solar Deployment and Subscriptions (As of January 2026)

The community solar market analysis presented here is based primarily on data collected through Sharing the Sun, an initiative of the National Community Solar Partnership+ (NCSP+). Sharing the Sun data collection and analysis are conducted by the National Laboratory of the Rockies (NLR) as part of its support for implementation of NCSP+. NLR first released a dataset of community solar projects in 2018 and updates it biannually. The January 2026 dataset, data collection methodology, and all the previous datasets are available from NLR's Data Catalog: https://data.nlr.gov/submissions/244. The dataset presents project-level information including location, capacity, operating utility, and year of interconnection. The dataset is created from multiple data sources such as utility data, public utility commissions, project developer websites, media releases, primary data collection by NLR, and data provided by developers under nondisclosure agreements. This presentation builds on a previous analysis of the community solar project dataset, Sharing the Sun: Community Solar Deployment and Subscriptions (as of June 2024). Dr. Gabriel Chan and his team at the University of Minnesota contribute to this effort. NCSP+ is led and funded by U.S. Department of Energy's Integrated Energy Systems Office (IESO).

14 SOLAR ENERGY↗

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY↗

Carbon Storage Site Mapping Inquiry Tool (MapIT)

To date, 48 projects, consisting of 139 wells, are currently under review with the Environmental Protection Agency’s (EPA) Underground Injection Control (UIC) Program for Class VI – wells used for geologic sequestration of carbon dioxide. The number of applications submitted is expected to increase in coming years with the increase of the 45Q tax credit available to projects that initiate construction prior to 2033. The amount of data collected to submit a Class VI permit is vast, and often disparate, coming from state, federal, and commercial entities, as well as field-specific data collected within an area of interest. When preparing for site selection and permitting, the initial aggregation of relevant public data can be time intensive. The Carbon Storage Site Mapping Inquiry tool (MapIT) was created to support and accelerate the discovery and accessibility of open-source data and information available across the USA. Data was aggregated and organized based on data types described within the EPA UIC Class VI permit documentation. The online tool enables users to explore hundreds of geospatial data layers and connect to additional external resources, leveraging API and REST services where possible to ensure updates to data in real time. MapIT enables users to explore state and federal data related to geologic, geophysical, structural, hydrologic, and contextual information. In addition to displaying spatial data and linking to external resources, MapIT leverages custom widgets to ensure that internal data and external data are discoverable and accessible. The widgets connect users to resources such as the USGS publications and the USGS Earthquake Catalog based on a user-defined location. This talk will describe data aggregation workflows, data types, data preparation, and tool development for MapIT. The Carbon Storage Site Mapping Inquiry Tool and underlying database are valuable, intuitive resources that empower government, academic, commercial and industry stakeholders to explore, analyze, and acquire carbon storage related data.

Morkner, Paige↗

New Particle Formation and Growth in the Houston Atmosphere During TRACER (Final Report)

From 2020-2025, researchers from UC Irvine, UC Riverside, and Colorado State University collaborated on a Department of Energy-funded project to understand how airborne particles form and grow in urban atmospheres, conducting an intensive field campaign in Houston, Texas during summer 2022. Using advanced instruments to measure gas-phase chemicals, particle composition, and a specialized chamber to study particle growth, the team discovered that sulfur-containing compounds from industrial and power plant emissions are the dominant driver of new particle formation in Houston, with particles typically forming locally in the city and growing as air moves away in the urban plume. The research revealed an important methodological insight: measurements from fixed ground stations can be misleading when interpreting how particles actually evolve as air masses move, which has significant implications for how scientists worldwide interpret atmospheric observations. These findings improve understanding of urban air quality and help reduce uncertainties in climate models, since these particles play critical roles in cloud formation and Earth's radiation balance, while also providing detailed information about ultrafine particle composition relevant to public health. The project trained three doctoral students, developed enhanced computer models for urban particle formation, and made all data publicly available through the DOE Atmospheric Radiation Measurement data archive for use by the broader scientific community.

54 ENVIRONMENTAL SCIENCES↗

Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns

The discovery and engineering of new enzymes is important across the bioeconomy, with diverse applications from foods to pharmaceuticals, sensors to agriculture. However, enzyme engineering, in particular machine learning-guided engineering, is hampered by a lack of data. Currently there exists no database designed to capture and interpret datasets created in this domain, nor are there easy analysis and visualisation tools. We developed the Enzyme Engineering Database to provide a centralized resource and an online analysis tool to consolidate sequence-function data from enzyme engineering campaigns, thereby making three contributions: (i) a database into which researchers can deposit public data, (ii) visualisation and analysis tools for protein engineers to analyse their own data or compare enzyme variants to other engineering campaigns, and (iii) a gold-standard dataset for benchmarking automated extraction along with the first large language model extraction pipeline specific for enzyme engineering campaigns. The Enzyme Engineering Database is accessible at http://enzengdb.org/.

Long, Yueming [California Institute of Technology ↗