Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “github”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

RHOD Site - NOAA PSL Wind Retrievals WINDoe / Derived Data

This dataset contains daily NetCDF files with horizontal wind profiles retrieved with the WINDoe retrieval (Gebauer and Bell 2024) at Rhode Island (RHOD). WINDoe retrievals datasets are also available at Nantucket Island (NANT, nant.windoe.z01.c1) and Block Island (BLOC, bloc.windoe.z01.c1). WINDoe is an optimal estimation algorithm to retrieve wind profiles combining multiple instruments. The code is available in this github repository (https://github.com/OAR-atmospheric-observations/WINDoe/tree/main) and the retrieval is described by Gebauer and Bell (2024). WINDoe allows combining the individual datasets and outputs into one profile taking into account the information and uncertainties of each dataset. The use of WINDoe minimizes data gaps and maximizes data availability, compared to using wind profiles from only one of the instruments. The regular height grid eases comparisons to numerical weather prediction models. Code modifications have been made that include reading in WFIP3 specific instruments, averaging Doppler lidar radial velocities at various azimuth angles to avoid overfitting, and allowing the user to define a height grid by the user in the vipfile. The instruments used as input to the retrieval are a radar wind profiler (low- and high resolution mode) providing data in and above the boundary layer, a scanning Doppler lidar usually providing data throughout the boundary layer, a profiling lidar providing data from 50 to 200 m at BLOC and NANT, and from 10 to 280 m at Rhode Island, and a surface tower (4 m at NANT and RHOD and 10 m at BLOC). From the scanning lidars, we used radial velocity measurements at 60 deg elevation angle at six different azimuth angles with a resolution of approximately 30 m along the line of sight and the lowest range gate at approximately 70 m. The wind profiles are retrieved with WINDoe up to 3.74 km with 10 m vertical resolution. The profiles are retrieved every 15 min at BLOC and NANT and every 60 min at RHOD.

17 WIND ENERGY↗

BLOC Site - NOAA PSL Wind Retrievals WINDoe / Derived Data

This dataset contains daily netcdf files with horizontal wind profiles retrieved with the WINDoe retrieval (Gebauer and Bell 2024) at Block Island (BLOC). WINDoe retrievals datasets are also available at Nantucket Island (NANT, nant.windoe.z01.c1) and Rhode Island (RHOD, rhod.windoe.z01.c1). WINDoe is an optimal estimation algorithm to retrieve wind profiles combining multiple instruments. The code is available in this github repository (https://github.com/OAR-atmospheric-observations/WINDoe/tree/main), and the retrieval is described by Gebauer and Bell (2024). WINDoe allows combining the individual datasets and outputs into one profile taking into account the information and uncertainties of each dataset. The use of WINDoe minimizes data gaps and maximizes data availability, compared to using wind profiles from only one of the instruments. The regular height grid eases comparisons to numerical weather prediction models. Code modifications have been made that include reading in WFIP3 specific instruments, averaging Doppler lidar radial velocities at various azimuth angles to avoid overfitting, and allowing the user to define a height grid by the user in the vipfile. The instruments used as input to the retrieval are a radar wind profiler (low- and high resolution mode) providing data in and above the boundary layer, a scanning Doppler lidar usually providing data throughout the boundary layer, a profiling lidar providing data from 50 to 200 m at BLOC and NANT, and from 10 to 280 m at Rhode Island, and a surface tower (4 m at NANT and RHOD and 10 m at BLOC). From the scanning lidars, we used radial velocity measurements at 60 deg elevation angle at six different azimuth angles with a resolution of approximately 30 m along the line of sight and the lowest range gate at approximately 70 m. The wind profiles are retrieved with WINDoe up to 3.74 km with 10 m vertical resolution. The profiles are retrieved every 15 min at BLOC and NANT and every 60 min at RHOD.

17 WIND ENERGY↗

NANT Site - NOAA PSL Wind Retrievals WINDoe / Derived Data

This dataset contains daily NetCDF files with horizontal wind profiles retrieved with the WINDoe retrieval (Gebauer and Bell 2024) at Nantucket Island (NANT). WINDoe retrievals datasets are also available at Block Island (BLOC, bloc.windoe.z01.c1) and Rhode Island (RHOD, rhod.windoe.z01.c1). WINDoe is an optimal estimation algorithm to retrieve wind profiles combining multiple instruments. The code is available in this github repository (https://github.com/OAR-atmospheric-observations/WINDoe/tree/main), and the retrieval is described by Gebauer and Bell (2024). WINDoe allows combining the individual datasets and outputs into one profile taking into account the information and uncertainties of each dataset. The use of WINDoe minimizes data gaps and maximizes data availability, compared to using wind profiles from only one of the instruments. The regular height grid eases comparisons to numerical weather prediction models. Code modifications have been made that include reading in WFIP3 specific instruments, averaging Doppler lidar radial velocities at various azimuth angles to avoid overfitting, and allowing the user to define a height grid by the user in the vipfile. The instruments used as input to the retrieval are a radar wind profiler (low- and high resolution mode) providing data in and above the boundary layer, a scanning Doppler lidar usually providing data throughout the boundary layer, a profiling lidar providing data from 50 to 200 m at BLOC and NANT, and from 10 to 280 m at Rhode Island, and a surface tower (4 m at NANT and RHOD and 10 m at BLOC). From the scanning lidars, we used radial velocity measurements at 60 deg elevation angle at six different azimuth angles with a resolution of approximately 30 m along the line of sight and the lowest range gate at approximately 70 m. The wind profiles are retrieved with WINDoe up to 3.74 km with 10 m vertical resolution. The profiles are retrieved every 15 min at BLOC and NANT and every 60 min at RHOD.

17 WIND ENERGY↗

NANT Site - ASSIST Thermodynamic Retrievals TROPoe v0.18 / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). This is a post-processed dataset and recommended for use. The profiles are retrieved every 10 minutes from instantaneous radiances observed with an Atmospheric Sounder Spectrometer by Infrared Spectral Technology (ASSIST, Michaud-Belleau et al. 2025) operated by NOAA Physical Sciences Laboratory (PSL) on Nantucket Island for WFIP3. The spectral bands used in the retrieval are in the wavenumber range from 612 - 905.4 cm-1 and are specified in Turner and Löhnert (2021). Additional input data in TROPoe are cloud base height from a collocated ceilometer operated by NOAA GML and temperature, water vapor mixing ratio, and pressure from a collocated surface tower operated by NOAA PSL. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Upton, NY. The TROPoe docker container (version 0.18) is available from Docker Hub at https://hub.docker.com/r/davidturner53/tropoe/tags, and the source code code is available in the GitHub repository https://github.com/OAR-atmospheric-observations/TROPoe.

17 WIND ENERGY↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Data for The utility of transfer learning to improve the performance of deep learning in axon segmentation

The utility of transfer learning to improve the performance of deep learning in axon segmentation Data Data: All the input and labeled volumes tf-logs: Tensorflow logs, view with command "tensorboard --logdir [name of folder]" Model Weights: model_weights: the argument list under variable combo indicate 1) no oversampling, 2) no rotation, 3) no learn scheduler, and 4) flipping on all three dimensions, and the additional values indicate 5) elastic deformation percentage, 6) rotate deformation percentage, 7) layer setting , 8) learning rate, and 9) training/validation/test data division suffix (leave '' if not using suffix). Results: Output from inference segment_total_results_validation_final: All validation results and calculations segment_total_results: All test results and calculations Authors The modified code was created for a paper by: Marjolein Oostrom, Michael A. Muniak, Rogene Eichler West, Sarah Akers, Paritosh Pande, Moses Obiri, Wei Wang, Kasey Bowyer, Zhuhao Wu, Lisa Bramer, Tianyi Mao, Bobbie Jo Webb-Robertson The work is adapted from Github TrailMap, which was created by Albert Pun and Drew Friedmann Acknowledgments MO, RMEW, SA, MO, LB, BJWR were supported by the Laboratory Directed Research and Development at Pacific Northwest National Laboratory (PNNL), a Department of Energy facility operated by Battelle under contract DE-AC05-76RLO01830. WW, KB, and ZW were supported in part by a NIH/BRAIN Initiative Grant RF1MH128969. MAM and TM were supported by two NIH/BRAIN Initiative Grants R01NS104944, RF1MH120119 and NIH R01NS081071. This research is affiliated with the Pacific northwest bioMedical Innovation Co-laboratory (PMedIC) collaboration between OHSU and PNNL.

Oostrom, Marjolein T↗

Nominal and adversarial synthetic PMU data for standard IEEE test systems

GridSTAGE (Spatio-Temporal Adversarial scenario GEneration) is a framework for the simulation of adversarial scenarios and the generation of multivariate spatio-temporal data in cyber-physical systems. GridSTAGE is developed based on Matlab and leverages Power System Toolbox (PST) where the evolution of the power network is governed by nonlinear differential equations. Using GridSTAGE, one can create several event scenarios that correspond to several operating states of the power network by enabling or disabling any of the following: faults, AGC control, PSS control, exciter control, load changes, generation changes, and different types of cyber-attacks. Standard IEEE bus system data is used to define the power system environment. GridSTAGE emulates the data from PMU and SCADA sensors. The rate of frequency and location of the sensors can be adjusted as well. Detailed instructions on generating data scenarios with different system topologies, attack characteristics, load characteristics, sensor configuration, control parameters are available in the Github repository - https://github.com/pnnl/GridSTAGE. There is no existing adversarial data-generation framework that can incorporate several attack characteristics and yield adversarial PMU data. The GridSTAGE framework currently supports simulation of False Data Injection attacks (such as a ramp, step, random, trapezoidal, multiplicative, replay, freezing) and Denial of Service attacks (such as time-delay, packet-loss) on PMU data. Furthermore, it supports generating spatio-temporal time-series data corresponding to several random load changes across the network or corresponding to several generation changes. A Koopman mode decomposition (KMD) based algorithm to detect and identify the false data attacks in real-time is proposed in https://ieeexplore.ieee.org/document/9303022. Machine learning-based predictive models are developed to capture the dynamics of the underlying power system with a high level of accuracy under various operating conditions for IEEE 68 bus system. The corresponding machine learning models are available at https://github.com/pnnl/grid_prediction.

99 GENERAL AND MISCELLANEOUS↗

DRAM example narrative

DRAM example narrative DRAM on KBase let's anyone run annotations using DRAM in the cloud. DRAM is an annotation tool that can annotate bacterial, archaeal and viral genomes and distills those annotatios into represetations of the functional genomic potential of those organisms. If you want to read more about DRAM you can check out the GitHub, wiki and journal article. DRAM annotate assemblies In KBase Assembly objects contain nucleotide sequences from genomes or metagenomes. DRAM can predict genes and annotate their function from KBase Assembly objects which may be microbial isolate genomes, metagenome assembled genomes or metagenomes. This is done with the Annotate and Distill Assemblies with DRAM app. This app can also anntoate AssemblySet objects which contain collection of Assembly objects. It also generates a Genome object and a GenomeSet object which can be used for further analysis with other KBase apps. The full annotations and other DRAM files are also available for download in the app.

59 BASIC BIOLOGICAL SCIENCES↗

PV Rooftop Database for Puerto Rico (PVRDB-PR)

The National Renewable Energy Laboratory's (NREL) PV Rooftop Database for Puerto Rico (PVRDB-PR) is a lidar-derived, geospatially-resolved dataset of suitable roof surfaces and their PV technical potential for virtually all buildings in Puerto Rico. The dataset can be downloaded at the AWS S3 explorer page. The GitHub documentation page provides a description of the dataset with methods and assumptions. The Puerto Rico Solar-For-All dataset provides Census Tract level estimates of residential low-to-moderate income (LMI) PV rooftop technical potential as well as solar electric bill savings potential for LMI communities at the municipality level.

Array↗

Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts (Sup3rCC)

The Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts (Sup3rCC) data is a collection of 4km hourly wind, solar, temperature, humidity, and pressure fields for the contiguous United States under various climate change scenarios. Sup3rCC is downscaled Global Climate Model (GCM) data. The downscaling process was performed using a generative machine learning approach called sup3r: Super-Resolution for Renewable Energy Resource Data (linked below as "Sup3r GitHub Repo"). The data includes both historical and future weather years, although the historical years represent the historical climate, not the actual historical weather that we experienced. You cannot use Sup3rCC data to study historical weather events, although other sup3r datasets may be intended for this. The Sup3rCC data is intended to help researchers study the impact of climate change on energy systems with high levels of wind and solar capacity. Please note that all climate change data is only a representation of the possible future climate and contains significant uncertainty. Analysis of multiple climate change scenarios and multiple climate models can help quantify this uncertainty.

Array↗

Linearized Distribution Optimal Power Flow for OEDI SI

This research is to meant to demonstrate the OEDI SI use case for distributed optimal power flow (DOPF). The goal was to formulate the optimal power flow problem in the distribution system for active and reactive power setpoints of PV systems using topology information and voltage measurements. The co-simulation runs every 15 minutes as outlined within the scenario file for the given feeder configuration. The linked GitHub repository includes five federates to achieve DOPF for the small, medium, large, and IEEE 123 feeder scenarios. We are using the OEDI SI framework, as well as the example feeder, sensor, recorder, and estimator federates provided in the example repository for OEDI SI. We also provide a runner script for switching between scenarios.

algorithm↗

Airfoil Computational Fluid Dynamics - 2k shapes, 25 AoA's, 3 Re numbers

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 1,830 airfoil shapes computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this dataset, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We sampled airfoil designs over both parameter spaces to explore the full range of possible shape variations. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, for or three different Reynolds' numbers (3M, 6M, and 9M), and for 25 angles of attack from -4 deg. to 20 deg. with 1 degree increments. Across all these various parameters, this dataset includes the results from over 250,000 CFD simulations. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy, in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_2k_data.ipynb.

2k↗

Airfoil Computational Fluid Dynamics - 9k shapes, 2 AoA's

This dataset contains aerodynamic quantities - including flow field values (momentum, energy, and vorticity) and summary values (coefficients of lift, drag, and momentum) - for 8,996 airfoil shapes, computed using the HAM2D CFD (computational fluid dynamics) model. The airfoil shapes were designed using the separable shape tensor parameterization that encodes two-dimensional shapes as elements of the Grassmann manifold. This data-driven approach learns two independent spaces of parameter from a collection of sample airfoils. The first captures large-scale, linear perturbations, and the second defines small-scale, higher-order perturbations. For this data, we used the G2Aero database of over 19,000 airfoil shapes to learn a parameter space that captured a wide array of shape characteristics. We fixed the linear deformations to be the mean over the database and sampled new shapes over a four-dimensional parameter space of higher-order perturbation. This sampling approaches allows for isolated analysis of non-linear airfoil shape deformations while holding other aspects (e.g., airfoil thickness) approximately constant. The aerodynamic quantities for the generated airfoil were obtained using the HAM2D code, which is a finite-volume Reynolds-averaged Navier-Stokes (RANS) flow solver. We employ a fifth-order WENO scheme for spatial reconstruction with Roe's flux difference scheme for inviscid flux and second-order central differencing for viscous flux. A preconditioned GMRES method is applied for implicit integration. The Spalart-Allmaras 1-eq turbulence model is used for the turbulence closure, and the Medida-Baeder 2-eq transition model is applied to account for the effects of laminar turbulent transition. The airfoil grid is generated with a total of 400 points on the airfoil surface, the initial wall-normal spacing of y+ = 1, and an outer boundary located at 300 chord lengths away from the wall. The CFD simulations are performed at a freestream Mach number of 0.1, Reynolds number of 9M, and at two angles of attack, 4 deg. and 12 deg. The simulations were performed using the Bridges-2 system at the Pittsburgh Supercomputing Center in February 2023 as part of the INTEGRATE project funded by the Advanced Research Projects Agency - Energy in the U.S. Department of Energy. The data was collected, reformatted, and preprocessed for this OEDI submission in July 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub Repository resource under explore_airfoil_9k_data.ipynb.

9k↗

Flow Redirection and Induction in Steady State (FLORIS) Wind Plant Power Production Data Sets

This dataset contains turbine- and plant-level power outputs for 252,500 cases of diverse wind plant layouts operating under a wide range of yawing and atmospheric conditions. The power outputs were computed using the Gaussian wake model in NREL's FLOw Redirection and Induction in Steady State (FLORIS) model, version 2.3.0. The 252,500 cases include 500 unique wind plants generated randomly by a specialized Plant Layout Generator (PLayGen) that samples randomized realizations of wind plant layouts from one of four canonical configurations: (i) cluster, (ii) single string, (iii) multiple string, (iv) parallel string. Other wind plant layout parameters were also randomly sampled, including the number of turbines (25-200) and the mean turbine spacing (3D-10D, where D denotes the turbine rotor diameter). For each layout, 500 different sets of atmospheric conditions were randomly sampled. These include wind speed in 0-25 m/s, wind direction in 0 deg.-360 deg., and turbulence intensity chosen from low (6%), medium (8%), and high (10%). For each atmospheric inflow scenario, the individual turbine yaw angles were randomly sampled from a one-sided truncated Gaussian on the interval 0 deg.-30 deg. oriented relative to wind inflow direction. This random data is supplemented with a collection of yaw-optimized samples where FLORIS was used to determine turbine yaw angles that maximize power production for the entire plant. To generate this data, a subset of cases were selected (50 atmospheric conditions from 50 layouts each for a total of additional 2,500 cases) for which FLORIS was re-run with wake steering control optimization. The IEA onshore reference turbine, which has a 130 m rotor diameter, a 110 m hub height, and a rated power capacity of 3.4 MW was used as the turbine for all simulations. The simulations were performed using NREL's Eagle high performance computing system in February 2021 as part of the Spatial Analysis for Wind Technology Development project funded by the U.S. Department of Energy Wind Energy Technologies Office. The data was collected, reformatted, and preprocessed for this OEDI submission in May 2023 under the Foundational AI for Wind Energy project funded by the U.S. Department of Energy Wind Energy Technologies Office. This dataset is intended to serve as a benchmark against which new artificial intelligence (AI) or machine learning (ML) tools may be tested. Baseline AI/ML methods for analyzing this dataset have been implemented, and a link to their repository containing those models has been provided. The .h5 data file structure can be found in the GitHub repository under explore_wind_plant_data_h5.ipynb.

AI↗

County-Level Hourly Renewable Capacity Factor Dataset for the ReEDS Model

This dataset contains hourly capacity factors for each renewable resource class and region (in this case, county). Technologies like large-scale utility PV (UPV), onshore wind, offshore wind, and concentrating solar power (CSP) are included. The dataset contains 7 years of hourly weather data (2007-2013) for different sites across the US and is used as one of the inputs to the ReEDS-2.0 model (see the "ReEDS 2.0 GitHub Repository" resource link below), developed by NREL. The weather profiles apply to any capacity that exists or is built in each region and class. This helps calculate the generation that can be provided using these resources. Open, reference, and limited are 3 scenarios based on land-use allowance, derived from the Renewable Energy Potential (reV) model developed by NREL, which helps generate supply curves for renewable technologies and assess the maximum potential of renewable resources in a designated area. Each zipped file in this dataset corresponds to a technology and contains the respective land-use scenario files required to run that technology in ReEDS. To use this dataset, download and place the extracted files in the locally cloned ReEDS repository inside one of the folders (inputs/variability/multi_year). After completing this copy, upon running the ReEDS model at the county-level spatial resolution for respective analysis purposes, the program will detect the presence of these files and will not fail.

Array↗

The Foundational Industry Energy Dataset: Unit-level Characterization and Derived Energy Estimates for Industrial Facilities in 2017

The Foundational Industry Energy Dataset (FIED) addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization by facility. Each facility is identified by a unique registryID, based on the U.S. Environmental Protection Agency (EPA) Facility Registry Service, and includes its coordinates and other geographic identifiers. Energy-using units are characterized by design capacity, as well as their estimated energy use, greenhouse gas emissions, and physical throughput using 2017 data from the EPA's National Emissions Inventory and Greenhouse Gas Reporting Program. An overview of the derivation methods is provided in a separate technical report which will be linked after publication. The Python code used to compile the dataset is available in a GitHub repository. An updated 2020 version is under development.

Array↗

The Journal of Open Source Software (JOSS): Bringing Open-Source Software Practices to the Scholarly Publishing Community for Authors, Reviewers, Editors, and Publishers

Open-source software (OSS) is a critical component of open science, but contributions to the OSS ecosystem are systematically undervalued in the current academic system. The Journal of Open Source Software (JOSS) contributes to addressing this by providing a venue (that is itself free, diamond open access, and all open-source, built in a layered structure using widely available elements/services of the scholarly publishing ecosystem) for publishing OSS, run in the style of OSS itself. A particularly distinctive element of JOSS is that it uses open peer review in a collaborative, iterative format, unlike most publishers. Additionally, all the components of the process—from the reviews to the papers to the software that is the subject of the papers to the software that the journal runs—are open. We describe JOSS’s history and its peer review process using an editorial bot, and we present statistics gathered from JOSS’s public review history on GitHub showing an increasing number of peer reviewed papers each year. We discuss the new JOSSCast and use it as a data source to understand reasons why interviewed authors decided to publish in JOSS. JOSS’s process differs significantly from traditional journals, which has impeded JOSS’s inclusion in indexing services such as Web of Science. In turn, this discourages researchers within certain academic systems, such as Italy’s, which emphasize the importance of Web of Science and/or Scopus indexing for grant applications and promotions. JOSS is a fully diamond open-access journal with a cost of around US$\$$5 per paper for the 401 papers published in 2023. The scalability of running JOSS with volunteers and financing JOSS with grants and donations is discussed.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Blending ToolKit: A simulation framework for evaluation of galaxy detection and deblending

We present an open source Python library for simulating overlapping (i.e., blended) images of galaxies and performing self-consistent comparisons of detection and deblending algorithms based on a suite of metrics. The package, named Blending Toolkit (BTK), serves as a modular, flexible, easy-to-install, and simple-to-use interface for exploring and analyzing systematic effects related to blended galaxies in cosmological surveys such as the Vera Rubin Observatory Legacy Survey of Space and Time (LSST). BTK has three main components: (1) a set of modules that perform fast image simulations of blended galaxies, using the open source image simulation package GalSim; (2) a module that standardizes the inputs and outputs of existing deblending algorithms; (3) a library of deblending metrics commonly defined in the galaxy deblending literature. In combination, these modules allow researchers to explore the impacts of galaxy blending in cosmological surveys. Additionally, BTK provides researchers who are developing a new deblending algorithm a framework to evaluate algorithm performance and make principled comparisons with existing deblenders. BTK includes a suite of tutorials and comprehensive documentation. The source code is publicly available on GitHub at https://github.com/LSSTDESC/BlendingToolKit.

79 ASTRONOMY AND ASTROPHYSICS↗