Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “github”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

3.0 - MOOSE: Enabling massively parallel multiphysics simulations

The development of MOOSE has kept accelerating since the last release, with over 2,100 pull requests merged over the last 30 months that involved nearly fifty contributors across close to a dozen institutions internationally. The growth in MOOSE's capabilities and downstream applications is reflected in the growth of the community. User support provided on the GitHub discussions forum has steadily increased to nearly 50 daily interactions. New simulation projects, notably to model advanced nuclear reactor and fusion devices, are driving a significant expansion of the capabilities. This paper reports on these developments, with several major released features, new physics modules, and key improvements to the user experience and simulation workflow.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

HARD: A performance portable radiation hydrodynamics code based on FleCSI framework

Hydrodynamics And Radiation Diffusion (HARD) is an open-source application for high-performance simulations of compressible hydrodynamics with radiation-diffusion coupling. Built on the FleCSI (Bergen et al., 2021 [1]) (Flexible Computational Science Infrastructure) framework, HARD expresses its computational units as tasks whose execution can be orchestrated by multiple back-end runtimes, including Legion (Bauer et al., 2012 [2]), MPI (Forum, 1994 [3]), and HPX (Kaiser et al., 2020 [4]). Node-level parallelism is handled through Kokkos (Edwards et al., 2014 [5]), providing a single-source, portable code base that runs efficiently on laptops, small homogeneous clusters, and the largest heterogeneous supercomputers currently available. To ensure scientific reliability, HARD includes a regression test suite that automatically reproduces canonical verification problems such as the Sod and LeBlanc shock tubes, and the Sedov blast wave, comparing numerical solutions against known analytical results. The project is distributed under an OSI-approved license, hosted on GitHub, and accompanied by reproducible build scripts and continuous integration workflows. This combination of performance portability, verification infrastructure, and community-focused development makes HARD a sustainable platform for advancing radiation hydrodynamics research across multiple domains.

97 MATHEMATICS AND COMPUTING↗

PyHydroGeophysX: An extensible open-source platform for integrating hydrological models with geophysical measurements

Hydrological models and geophysical measurements are widely used tools for understanding subsurface hydrological processes relevant to water resource management, yet they typically remain disconnected due to technical barriers. We present PyHydroGeophysX, an open-source Python platform bridging this gap by providing standardized interfaces between hydrological modeling software (MODFLOW, ParFlow) and geophysical simulation tools (PyGIMLi, SimPEG). The platform implements bidirectional workflows: translating hydrological outputs into simulated geophysical responses through petrophysical models, and extracting hydrological information from geophysical inversions. Key features include bidirectional workflow modules, configurable petrophysical models, time-lapse inversion with temporal regularization, parallel computing, and mesh utilities for property transfer between geophysical and hydrological grids. The modular architecture of PyHydroGeophysX enables researchers to incorporate additional models and methods, fostering broader adoption of integrated hydrogeophysical approaches. The software is freely available on GitHub and is intended for researchers and practitioners working at the intersection of hydrology and geophysics.

Hydrogeophysics↗

GBOpt: Grain boundary structure optimization using Monte Carlo and evolutionary algorithms

Polycrystalline materials are made of many small crystals separated by grain boundaries (GBs), whose atomic structure strongly influences material properties. Because the structure of a GB determines its properties, the optimal structure must be known in order to determine those impacts. There are many ways of placing atoms in the GB region, but the optimal structure is defined as the one that gives the lowest value of a target property (typically energy). GB structure optimization has been successfully demonstrated using stochastic and evolutionary methods, but no reusable, community-maintained open-source workflow has been developed. GBOpt (Grain Boundary Optimization) is an open-source Python package that creates that workflow, where we have presently implemented two approaches: Markov Chain Monte Carlo, and genetic algorithm based on elite selection. We demonstrate this capability by successfully reproducing the known optimal structures of a specific GB in two materials, and point interested readers to the GitHub repository for additional examples, including optimization for different properties. Both of the implemented approaches recovered the known structures, with the genetic algorithm approach finding the optimal structure faster on average.

99 - GENERAL AND MISCELLANEOUS↗

Benchmark probabilistic solar forecasts: Characteristics and recommendations

We illustrate and compare commonly used benchmark, or reference, methods for probabilistic solar forecasting that researchers use to measure the performance of their proposed techniques. A thorough review of the literature indicates wide variation in the benchmarks implemented in probabilistic solar forecast studies. To promote consistent and sensible methodological comparisons, we implement and compare ten variants from six common benchmark classes at two temporal scales: intra-hourly forecasts and hourly resolution forecasts. Using open-source Surface Radiation Budget Network (SURFRAD) data from 2018, these benchmark methods are compared using proper probabilistic metrics and common diagnostic tools. Practical implementation issues, such as the impact of missing data and applicability for operational forecasting, are also discussed. Furthermore, we make recommendations for practitioners on the appropriate selection of benchmark methods to properly showcase state-of-the-art improvements in forecast reliability and sharpness. All code and open-source data are available on Github for reproducibility and for other researchers to apply the same benchmark methods to their own data.

14 SOLAR ENERGY↗

pvlib iotools—Open-source Python functions for seamless access to solar irradiance data

Access to accurate solar resource data is critical for numerous applications, including estimating the yield of solar energy systems, developing radiation models, and validating irradiance datasets. However, lack of standardization in data formats and access interfaces across providers constitutes a major barrier to entry for new users. pvlib python’s iotools subpackage aims to solve this issue by providing standardized Python functions for reading local files and retrieving data from external providers. All functions follow a uniform pattern and return convenient data outputs, allowing users to seamlessly switch between data providers and explore alternative datasets. The pvlib package is community-developed on GitHub: https://github.com/pvlib/pvlib-python. As of pvlib python version 0.9.5, the iotools subpackage supports 12 different datasets, including ground measurement, reanalysis, and satellite-derived irradiance data. The supported ground measurement networks include the Baseline Surface Radiation Network (BSRN), NREL MIDC, SRML, SOLRAD, SURFRAD, and the US Climate Reference Network (CRN). Additionally, satellite-derived and reanalysis irradiance data from the following sources are supported: PVGIS (SARAH & ERA5), NSRDB PSM3, and CAMS Radiation Service (including McClear clear-sky irradiance).

14 SOLAR ENERGY↗

Uncertainty quantification in machine learning for engineering design and health prognostics: A tutorial

On top of machine learning (ML) models, uncertainty quantification (UQ) functions as an essential layer of safety assurance that could lead to more principled decision making by enabling sound risk assessment and management. The safety and reliability improvement of ML models empowered by UQ has the potential to significantly facilitate the broad adoption of ML solutions in high-stakes decision settings, such as healthcare, manufacturing, and aviation, to name a few. In this tutorial, we aim to provide a holistic lens on emerging UQ methods for ML models with a particular focus on neural networks and the applications of these UQ methods in tackling engineering design as well as prognostics and health management problems. Towards this goal, we start with a comprehensive classification of uncertainty types, sources, and causes pertaining to UQ of ML models. Next, we provide a tutorial-style description of several state-of-the-art UQ methods: Gaussian process regression, Bayesian neural network, neural network ensemble, and deterministic UQ methods focusing on spectral-normalized neural Gaussian process. Established upon the mathematical formulations, we subsequently examine the soundness of these UQ methods quantitatively and qualitatively (by a toy regression example) to examine their strengths and shortcomings from different dimensions. Then, we review quantitative metrics commonly used to assess the quality of predictive uncertainty in classification and regression problems. Afterward, we discuss the increasingly important role of UQ of ML models in solving challenging problems in engineering design and health prognostics. In conclusion, two case studies with source codes available on GitHub are used to demonstrate these UQ methods and compare their performance in the life prediction of lithium-ion batteries at the early stage (case study 1) and the remaining useful life prediction of turbofan engines (case study 2).

97 MATHEMATICS AND COMPUTING↗

Improving Enzyme Optimum Temperature Prediction with Resampling Strategies and Ensemble Learning

Accurate prediction of the optimal catalytic temperature ( T opt ) of enzymes is vital in biotechnology, as enzymes with high T opt values are desired for enhanced reaction rates. Recently, a machine learning method (temperature optima for microorganisms and enzymes, TOME) for predicting T opt was developed. TOME was trained on a normally distributed data set with a median T opt of 37 °C and less than 5% of T opt values above 85 °C, limiting the method’s predictive capabilities for thermostable enzymes. Due to the distribution of the training data, the mean squared error on T opt values greater than 85 °C is nearly an order of magnitude higher than the error on values between 30 and 50 °C. Here, we apply ensemble learning and resampling strategies that tackle the data imbalance to significantly decrease the error on high T opt values (>85 °C) by 60% and increase the overall R 2 value from 0.527 to 0.632. The revised method, temperature optima for enzymes with resampling (TOMER), and the resampling strategies applied in this work are freely available to other researchers as Python packages on GitHub.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NEXTorch: A Design and Bayesian Optimization Toolkit for Chemical Sciences and Engineering

Automation and optimization of chemical systems require well-informed decisions on what experiments to run to reduce time, materials, and/or computations. Data-driven active learning algorithms have emerged as valuable tools to solve such tasks. Bayesian optimization, a sequential global optimization approach, is a popular active-learning framework. Past studies have demonstrated its efficiency in solving chemistry and engineering problems. Here we introduce NEXTorch, a library in Python/PyTorch, to facilitate laboratory or computational design using Bayesian optimization. NEXTorch offers fast predictive modeling, flexible optimization loops, visualization capabilities, easy interfacing with legacy software, and multiple types of parameters and data type conversions. It provides GPU acceleration, parallelization, and state-of-the-art Bayesian optimization algorithms and supports both automated an d human-in-the-loop optimization. The comprehensive online documentation introduces Bayesian optimization theory and several examples from catalyst synthesis, reaction condition optimization, parameter estimation, and reactor geometry optimization. NEXTorch is open-source and available on GitHub

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Adaptive Ensemble Refinement of Protein Structures in High Resolution Electron Microscopy Density Maps with Radical Augmented Molecular Dynamics Flexible Fitting

Recent advances in cryo-electron microscopy (cryo-EM) have enabled modeling macromolecular complexes that are essential components of the cellular machinery. The density maps derived from cryo-EM experiments are often integrated with manual, knowledge or artificial intelligence driven, and physics-guided computational methods to build, fit, and refine molecular structures. Going beyond a single stationary- structure determination scheme, it is becoming more common to interpret the experimental data with an ensemble of models, which contributes to an average observation. Hence, there is a need to decide on the quality of an ensemble of protein structures on-the-fly, while refining them against the density maps. Here, we introduce such an adaptive decision making scheme during the molecular dynamics flexible fitting (MDFF) of biomolecules. Using RADICAL-Cybertools, and the new RADICAL augmented MDFF implementation (R-MDFF) is examined in high-performance computing environments for refinement of two protein systems, Adenylate Kinase and Carbon Monoxide Dehydrogenase. For the test cases, use of multiple replicas in flexible fitting with adaptive decision making in R-MDFF improves the overall correlation to the density by 40% relative to the refinements of the brute-force MDFF. The improvements are particularly significant at high, 2 - 3 Å, map resolutions. More importantly, the ensemble model captures key features of biologically relevant molecular dynamics that is inaccessible to a single-model interpretation. Finally, the pipeline is applicable to systems of growing sizes, which is demonstrated using ensemble refinement of capsid proteins from Chimpanzee adenovirus. The overhead for decision making remaining low and robust to computing environments. The software is publicly available on GitHub and includes a short user guide to install the R-MDFF on different computing environments, from local Linux based workstations to High Performance Computing (HPC) environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NWChem: Recent and Ongoing Developments

In this paper we summarize developments in the NWChem computational chemistry suite since the last major release (NWChem 7.0). Specifically, we focus on functionalities, along with input blocks, that are currently accessible in the current stable release (NWChem 7.2) and master branches, interfaces to quantum computing simulators, interfaces to external libraries, the NWChem GitHub repository, and containerization of NWChem executable images. In conclusion, some of the ongoing developments that will be available in the near future are also discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

leapR: An R Package for Multiomic Pathway Analysis

A generalized goal of many high-throughput data studies is to identify functional mecha-nisms that underlie observed biological phenomena, whether disease outcomes or metabolic out-put. Increasingly, studies that rely on multiple sources of high-throughput data (genomic, tran-scriptomic, proteomic, metabolomic) are faced with a challenge of utilizing the data in a way that maximizes utility. However, methods for integration of multiple forms of molecular data into a biolog-ically coherent frameworks are needed. Furthermore, we have developed a framework to assess biological pathway activity that relates to phenotypic outcome using multi-source data. Availability and implementation: The leapR package with user manual and example workflow is available for download from GitHub (https://github.com/biodataganache/leapR).

59 BASIC BIOLOGICAL SCIENCES↗

Structure Prediction of Ionic Epitaxial Interfaces with Ogre Demonstrated for Colloidal Heterostructures of Lead Halide Perovskites

Colloidal epitaxial heterostructures are nanoparticles composed of two different materials connected at an interface, which can exhibit properties different from those of their individual components. Combining dissimilar materials offers exciting opportunities to create a wide variety of functional heterostructures. However, assessing structural compatibility–the main prerequisite for epitaxial growth–is challenging when pairing complex materials with different lattice parameters and crystal structures. This complicates both the selection of target heterostructures for synthesis and the assignment of interface models when new heterostructures are obtained. Here, we demonstrate Ogre as a powerful tool to accelerate the design and characterization of colloidal heterostructures. To this end, we implemented developments tailored for the high-efficiency prediction of epitaxial interfaces between ionic/polar materials, which encompass most colloidal semiconductors. These include the use of pre-screening candidate models based on charge balance at the interface and the use of a classical potential for fast energy evaluations, with parameters automatically calculated based on the input bulk structures. These developments are validated for perovskite-based CsPbBr 3 /Pb 4 S 3 Br 2 heterostructures, where Ogre produces interface models in excellent agreement with density functional theory and experiments. Furthermore, we use Ogre to rationalize the templating effect of CsPbCl 3 on the growth of lead sulfochlorides, where perovskite seeds induce the formation of Pb 4 S 3 Cl 2 rather than Pb 3 S 2 Cl 2 due to better epitaxial compatibility. Finally, combining Ogre simulations with experimental data enables us to unravel the structure and composition of the hitherto unsolved CsPbBr 3 /Bi x Pb y S z interface, and to assign a structure to several other reported metal halide- and oxide-based interfaces. The Ogre package is available on GitHub or via the OgreInterface desktop application, available for Windows, Linux, and Mac.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The United States COVID-19 Forecast Hub dataset

Academic researchers, government agencies, industry groups, and individuals have produced forecasts at an unprecedented scale during the COVID-19 pandemic. To leverage these forecasts, the United States Centers for Disease Control and Prevention (CDC) partnered with an academic research lab at the University of Massachusetts Amherst to create the US COVID-19 Forecast Hub. Launched in April 2020, the Forecast Hub is a dataset with point and probabilistic forecasts of incident cases, incident hospitalizations, incident deaths, and cumulative deaths due to COVID-19 at county, state, and national, levels in the United States. Included forecasts represent a variety of modeling approaches, data sources, and assumptions regarding the spread of COVID-19. The goal of this dataset is to establish a standardized and comparable set of short-term forecasts from modeling teams. These data can be used to develop ensemble models, communicate forecasts to the public, create visualizations, compare models, and inform policies regarding COVID-19 mitigation. These open-source data are available via download from GitHub, through an online API, and through R packages.

60 APPLIED LIFE SCIENCES↗

Geometry-complete diffusion for 3D molecule generation and optimization

Abstract Generative deep learning methods have recently been proposed for generating 3D molecules using equivariant graph neural networks (GNNs) within a denoising diffusion framework. However, such methods are unable to learn important geometric properties of 3D molecules, as they adopt molecule-agnostic and non-geometric GNNs as their 3D graph denoising networks, which notably hinders their ability to generate valid large 3D molecules. In this work, we address these gaps by introducing the Geometry-Complete Diffusion Model (GCDM) for 3D molecule generation, which outperforms existing 3D molecular diffusion models by significant margins across conditional and unconditional settings for the QM9 dataset and the larger GEOM-Drugs dataset, respectively. Importantly, we demonstrate that GCDM’s generative denoising process enables the model to generate a significant proportion of valid and energetically-stable large molecules at the scale of GEOM-Drugs, whereas previous methods fail to do so with the features they learn. Additionally, we show that extensions of GCDM can not only effectively design 3D molecules for specific protein pockets but can be repurposed to consistently optimize the geometry and chemical composition of existing 3D molecules for molecular stability and property specificity, demonstrating new versatility of molecular diffusion models. Code and data are freely available on GitHub .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Building workflows for an interactive human-in-the-loop automated experiment (hAE) in STEM-EELS

Exploring the structural, chemical, and physical properties of matter on the nano- and atomic scales has become possible with the recent advances in aberration-corrected electron energy-loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). However, the current paradigm of STEM-EELS relies on the classical rectangular grid sampling, in which all surface regions are assumed to be of equal a priori interest. However, this is typically not the case for real-world scenarios, where phenomena of interest are concentrated in a small number of spatial locations, such as interfaces, structural and topological defects, and multi-phase inclusions. One of the foundational problems is the discovery of nanometer- or atomic-scale structures having specific signatures in EELS spectra. Herein, we systematically explore the hyperparameters controlling deep kernel learning (DKL) discovery workflows for STEM-EELS and identify the role of the local structural descriptors and acquisition functions in experiment progression. In agreement with the actual experiment, we observe that for certain parameter combinations the experiment path can be trapped in the local minima. We demonstrate the approaches for monitoring the automated experiment in the real and feature space of the system and knowledge acquisition of the DKL model. Based on these, we construct intervention strategies defining the human-in-the-loop automated experiment (hAE). This approach can be further extended to other techniques including 4D STEM and other forms of spectroscopic imaging. The hAE library is available on Github at https://github.com/utkarshp1161/hAE/tree/main/hAE.

Pratiush, Utkarsh [Univ. of Tennessee, Knoxville, ↗

Publication of the Belle II Software

The Belle II software was developed by a few hundred individual contributors over several years. Following the rising desire of making it publicly available, the collaboration established open source software policies and procedures. The political and technical challenges and their solutions at Belle II are discussed in this article. With the publication of the Belle II software, basf2, on GitHub and Zenodo in 2021 an important milestone towards open science was reached.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Consistent and reproducible computation of the glass transition temperature from molecular dynamics simulations

In many fields, from semiconductors for opto-electronic applications to ionic liquids (ILs) for separations, the glass transition temperature (Tg) of a material is a useful gauge for its potential use in practical settings. As a result, there is a great deal of interest in predicting Tg using molecular simulations. However, the uncertainty and variation in the trend shift method, a common approach in simulations to predict Tg, can be high. This is due to the need for human intervention in defining a fitting range for linear fits of density with temperature assumed for the liquid and glass phases across the simulated cooling. The definition of such fitting ranges then defines the estimate for the Tg as the intersection of linear fits. We eliminate this need for human intervention by leveraging the Shapiro–Wilk normality test and proposing an algorithm to define the fitting ranges and, consequently, Tg. Through this integration, we incorporate into our automated methodology that residuals must be normally distributed around zero for any fit, a requirement that must be met for any regression problem. Consequently, fitting ranges for realizing linear fits for each phase are statistically defined rather than visually inferred, obtaining an estimate for Tg without any human intervention. The method is also capable of finding multiple linear regimes across density vs temperature curves. We compare the predictions of our proposed method across multiple IL and semiconductor molecular dynamics simulation results from the literature and compare other proposed methods for automatically detecting Tg from density–temperature data. We believe that our proposed method would allow for more consistent predictions of Tg. We make this methodology available and open source through GitHub.

Chemistry↗