Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “github”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Building workflows for an interactive human-in-the-loop automated experiment (hAE) in STEM-EELS

Exploring the structural, chemical, and physical properties of matter on the nano- and atomic scales has become possible with the recent advances in aberration-corrected electron energy-loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). However, the current paradigm of STEM-EELS relies on the classical rectangular grid sampling, in which all surface regions are assumed to be of equal a priori interest. However, this is typically not the case for real-world scenarios, where phenomena of interest are concentrated in a small number of spatial locations, such as interfaces, structural and topological defects, and multi-phase inclusions. One of the foundational problems is the discovery of nanometer- or atomic-scale structures having specific signatures in EELS spectra. Herein, we systematically explore the hyperparameters controlling deep kernel learning (DKL) discovery workflows for STEM-EELS and identify the role of the local structural descriptors and acquisition functions in experiment progression. In agreement with the actual experiment, we observe that for certain parameter combinations the experiment path can be trapped in the local minima. We demonstrate the approaches for monitoring the automated experiment in the real and feature space of the system and knowledge acquisition of the DKL model. Based on these, we construct intervention strategies defining the human-in-the-loop automated experiment (hAE). This approach can be further extended to other techniques including 4D STEM and other forms of spectroscopic imaging. The hAE library is available on Github at https://github.com/utkarshp1161/hAE/tree/main/hAE.

Pratiush, Utkarsh [Univ. of Tennessee, Knoxville, ↗

Publication of the Belle II Software

The Belle II software was developed by a few hundred individual contributors over several years. Following the rising desire of making it publicly available, the collaboration established open source software policies and procedures. The political and technical challenges and their solutions at Belle II are discussed in this article. With the publication of the Belle II software, basf2, on GitHub and Zenodo in 2021 an important milestone towards open science was reached.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Consistent and reproducible computation of the glass transition temperature from molecular dynamics simulations

In many fields, from semiconductors for opto-electronic applications to ionic liquids (ILs) for separations, the glass transition temperature (Tg) of a material is a useful gauge for its potential use in practical settings. As a result, there is a great deal of interest in predicting Tg using molecular simulations. However, the uncertainty and variation in the trend shift method, a common approach in simulations to predict Tg, can be high. This is due to the need for human intervention in defining a fitting range for linear fits of density with temperature assumed for the liquid and glass phases across the simulated cooling. The definition of such fitting ranges then defines the estimate for the Tg as the intersection of linear fits. We eliminate this need for human intervention by leveraging the Shapiro–Wilk normality test and proposing an algorithm to define the fitting ranges and, consequently, Tg. Through this integration, we incorporate into our automated methodology that residuals must be normally distributed around zero for any fit, a requirement that must be met for any regression problem. Consequently, fitting ranges for realizing linear fits for each phase are statistically defined rather than visually inferred, obtaining an estimate for Tg without any human intervention. The method is also capable of finding multiple linear regimes across density vs temperature curves. We compare the predictions of our proposed method across multiple IL and semiconductor molecular dynamics simulation results from the literature and compare other proposed methods for automatically detecting Tg from density–temperature data. We believe that our proposed method would allow for more consistent predictions of Tg. We make this methodology available and open source through GitHub.

Chemistry↗

Computational toolkit for predicting thickness of 2D materials using machine learning and autogenerated dataset by large language model

The thickness of 2D materials not only plays a crucial role in determining the performance of nanoelectronic and optoelectronic devices but also introduces complexities in predicting volume-dependent properties, such as energy storage capacity, due to the intrinsic vacuum within these materials. Although a plethora of experimental techniques, including but not limited to optical contrast, Raman spectroscopy, nonlinear optical spectroscopy, near-field optical imaging, and hyperspectral imaging, facilitate the measurement of 2D material thickness, comprehensive data for many materials remain elusive. Over the past decade, the exponential proliferation of 2D materials and their heterostructures has outstripped the capabilities of conventional experimental and computational approaches. In this evolving landscape, machine learning (ML) has emerged as an indispensable tool, offering a scalable approach to augment these traditional methodologies. Addressing the critical gap, we introduce THICK2D—Thickness Hierarchy Inference and Calculation Kit for 2D Materials. This Python-based computational framework harnesses an autogenerated thickness database, developed using large language models, and advanced ML algorithms to facilitate the rapid and scalable estimation of material thickness, relying solely on crystallographic data. To demonstrate the utility and robustness of THICK2D, we successfully used the toolkit to predict the thickness of more than 8000 2D-based materials, sourced from two extensive 2D materials databases. THICK2D is disseminated as an open-source utility, accessible on GitHub at https://github.com/gmp007/THICK2D, and archived on Zenodo at https://10.5281/zenodo.11216648.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Continuous Integration, In-Code Documentation, and Automation for Nuclear Quality Assurance Conformance

The Multiphysics Object Oriented Simulation Environment (MOOSE) is an open-source, finite element framework for solving highly coupled sets of nonlinear equations. The development of the framework and applications occurs concurrently using an agile, continuous-integration software package. Included in the framework is an in-code, extensible documentation system. Using these two tools in union with the repository management tools GitHub and GitLab, a software quality plan was created and followed such that MOOSE and a MOOSE-based application (BISON) have been shown to meet the American Society of Mechanical Engineers’ Nuclear Quality Assurance-1 standard. The approach relies heavily on automation for both testing and documentation. The resulting effort demonstrates that a rigorous software quality plan may be implemented that incurs a minimal impact on day-to-day development of the software, satisfying the stringent guidelines necessary to operate the software in a safety function within a nuclear facility.

97 MATHEMATICS AND COMPUTING↗

Versatile recognition of graphene layers from optical images under controlled illumination through green channel correlation method

In this study, a simple yet versatile method is proposed for identifying the number of exfoliated graphene layers transferred on an oxide substrate from optical images, utilizing a limited number of input images for training, paired with a more traditional number of a few thousand well-published Github images for testing and predicting. Two thresholding approaches, namely the standard deviation-based approach and the linear regression-based approach, were employed in this study. The method specifically leverages the red, green, and blue color channels of image pixels and creates a correlation between the green channel of the background and the green channel of the various layers of graphene. This method proves to be a feasible alternative to deep learning-based graphene recognition and traditional microscopic analysis. The proposed methodology performs well under conditions where the effect of surrounding light on the graphene-on-oxide sample is minimum and allows rapid identification of the various graphene layers. Here, the study additionally addresses the functionality of the proposed methodology with nonhomogeneous lighting conditions, showcasing successful prediction of graphene layers from images that are lower in quality compared to typically published in literature. In all, the proposed methodology opens up the possibility for the non-destructive identification of graphene layers from optical images by utilizing a new and versatile method that is quick, inexpensive, and works well with fewer images that are not necessarily of high quality.

36 MATERIALS SCIENCE↗

Hazma meets HERWIG4DM: precision gamma-ray, neutrino, and positron spectra for light dark matter

Herein we present a new open-source package, Hazma 2, that computes accurate spectra relevant for indirect dark matter searches for photon, neutrino, and positron production from vector-mediated dark matter annihilation and for spin-one dark matter decay. The tool bridges across the regimes of validity of two state of the art codes: Hazma 1, which provides an accurate description below hadronic resonances up to center-of-mass energies around 250 MeV, and Herwig4DM, which is based on vector meson dominance and measured form factors, and accurate well into the few GeV range. The applicability of the combined code extends to approximately 1.5 GeV, above which the number of final state hadrons off of which we individually compute the photon, neutrino, and positron yield grows exceedingly rapidly. We provide example branching ratios, particle spectra and conservative observational constraints from existing gamma-ray data for the well-motivated cases of decaying dark photon dark matter and vector-mediated fermionic dark matter annihilation. Finally, we compare our results to other existing codes at the boundaries of their respective ranges of applicability. Hazma 2 is freely available on GitHub athttps://github.com/LoganAMorrison/Hazma.

79 ASTRONOMY AND ASTROPHYSICS↗

Wavelet flow for extragalactic foreground simulations

Extragalactic foregrounds in cosmic microwave background (CMB) observations are both a source of cosmological and astrophysical information and a nuisance to the CMB. Effective field-level modeling that captures their non-Gaussian statistical distributions is increasingly important for optimal information extraction, particularly given the low-noise observations from current and upcoming experiments. Here, we explore the use of Wavelet Flow (WF) models to tackle the novel task of modeling the field-level probability distributions of multi-component CMB secondaries and foregrounds. Specifically, we jointly train correlated CMB lensing convergence (κ) and cosmic infrared background (CIB) maps with a WF model and obtain a network that statistically recovers the input to high accuracy — the trained network generates samples of κ and CIB fields whose average power spectra are within a few percent of the inputs across all scales, and whose Minkowski functionals are similarly accurate compared to the inputs. Leveraging the multiscale architecture of these models, we fine-tune both the model parameters and the priors at each scale independently, optimizing performance across different resolutions. These results demonstrate that WF models can accurately simulate correlated components of CMB secondaries, supporting improved analysis of cosmological data. Our code and trained models can be found on this GitHub repo.

cosmological simulations↗

A limit on the total lepton number in the Universe from BBN and the CMB

At temperatures below the QCD phase transition, any substantial lepton number in the Universe can only be present within the neutrino sector. In this work, we systematically explore the impact of a non-vanishing lepton number on Big Bang Nucleosynthesis (BBN) and the Cosmic Microwave Background (CMB). Relying on our recently developed framework based on momentum averaged quantum kinetic equations for the neutrino density matrix, we solve the full BBN reaction network to obtain the abundances of primordial elements. We find that the maximal primordial total lepton number L allowed by BBN and the CMB is -0.12 (-0.10) ≤ L ≤ 0.13 (0.12) for NH (IH), while specific flavor directions can be even more constrained. This bound is complementary to the limits obtained from avoiding baryon overproduction through sphaleron processes at the electroweak phase transition since, although numerically weaker, it applies at lower temperatures and is obtained completely independently. We publicly release the C++ code COFLASY-C on GitHub (https://github.com/mariofnavarro/COFLASY/tree/COFLASY-C) which solves for the evolution of the neutrino quantum kinetic equations numerically.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Vertex-finding and reconstruction of contained two-track neutrino events in the MicroBooNE detector

In this work, we describe algorithms developed to isolate and accurately reconstruct two-track events that are contained within the MicroBooNE detector. This method is optimized to reconstruct two tracks of lengths longer than 5cm. This code has applications to searches for neutrino oscillations and measurements of cross sections using quasi-elastic-like charged current events. The algorithms we discuss will be applicable to all detectors running in Fermilab's Short Baseline Neutrino program (SBN), and to any future liquid argon time projection chamber (LArTPC) experiment with beam energies ~ 1 GeV. The algorithms are publicly available on a GITHUB repository. This reconstruction offers a complementary and independent alternative to the Pandora reconstruction package currently in use in LArTPC experiments, and provides similar reconstruction performance for two-track events.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

FINETUNA: fine-tuning accelerated molecular simulations

Abstract Progress towards the energy breakthroughs needed to combat climate change can be significantly accelerated through the efficient simulation of atomistic systems. However, simulation techniques based on first principles, such as density functional theory (DFT), are limited in their practical use due to their high computational expense. Machine learning approaches have the potential to approximate DFT in a computationally efficient manner, which could dramatically increase the impact of computational simulations on real-world problems. However, they are limited by their accuracy and the cost of generating labeled data. Here, we present an online active learning framework for accelerating the simulation of atomic systems efficiently and accurately by incorporating prior physical information learned by large-scale pre-trained graph neural network models from the Open Catalyst Project. Accelerating these simulations enables useful data to be generated more cheaply, allowing better models to be trained and more atomistic systems to be screened. We also present a method of comparing local optimization techniques on the basis of both their speed and accuracy. Experiments on 30 benchmark adsorbate-catalyst systems show that our method of transfer learning to incorporate prior information from pre-trained models accelerates simulations by reducing the number of DFT calculations by 91%, while meeting an accuracy threshold of 0.02 eV 93% of the time. Finally, we demonstrate a technique for leveraging the interactive functionality built in to Vienna ab initio Simulation Package (VASP) to efficiently compute single point calculations within our online active learning framework without the significant startup costs. This allows VASP to work in tandem with our framework while requiring 75% fewer self-consistent cycles than conventional single point calculations. The online active learning implementation, and examples using the VASP interactive code, are available in the open source FINETUNA package on Github.

97 MATHEMATICS AND COMPUTING↗

A robust estimator of mutual information for deep learning interpretability

Abstract We develop the use of mutual information (MI), a well-established metric in information theory, to interpret the inner workings of deep learning (DL) models. To accurately estimate MI from a finite number of samples, we present GMM-MI (pronounced ‘Jimmie’), an algorithm based on Gaussian mixture models that can be applied to both discrete and continuous settings. GMM-MI is computationally efficient, robust to the choice of hyperparameters and provides the uncertainty on the MI estimate due to the finite sample size. We extensively validate GMM-MI on toy data for which the ground truth MI is known, comparing its performance against established MI estimators. We then demonstrate the use of our MI estimator in the context of representation learning, working with synthetic data and physical datasets describing highly non-linear processes. We train DL models to encode high-dimensional data within a meaningful compressed (latent) representation, and use GMM-MI to quantify both the level of disentanglement between the latent variables, and their association with relevant physical quantities, thus unlocking the interpretability of the latent representation. We make GMM-MI publicly available in this GitHub repository.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Mic-hackathon 2024: hackathon on machine learning for electron and scanning probe microscopy

Microscopy is one of the primary sources of information on materials structure and functionality at the nanometer and atomic scales. The data generated through microscopy is often contained in well-structured datasets, enriched with extensive metadata and sample histories, although not always with the same level of detail or storage format. The broad incorporation of data management plans by major funding agencies ensures the preservation and accessibility of this data. However, deriving insights from these rich datasets remains challenging due to the lack of established code ecosystems, standardized benchmarks, and integration strategies. Correspondingly, the efficiency of data usage is very low, and time expenditures at the analysis stage are enormous. In addition to post-acquisition data analysis, the emergence of application programming interfaces by major microscope manufacturers now creates opportunities for real-time ML-based data analytics to enable automated decision making, and particularly ML-agent controlled real-time microscope operation. Despite these opportunities, there is a significant gap in integrating the ML community with the broader microscopy community, limiting the value that these methods bring to physics and materials discovery and materials optimization. Hackathons address these challenges by fostering collaboration between ML experts and microscopy professionals, encouraging the development of innovative solutions that leverage ML for microscopy and preparing the workforce of the future both for microscopy-intensive domains areas, instrument manufacturers, and ML scientists interested in real world applications for fundamental research, materials optimization, and manufacturing. The hackathon generated benchmark datasets and digital twins of microscopes that further contribute to the development of the field and establish data analysis ecosystems. All the codes can be found at GitHub(https://github.com/KalininGroup/Mic-hackathon-2024-codes-publication/tree/1.0.0.1) and Zenodo (https://zenodo.org/records/15579940).

97 MATHEMATICS AND COMPUTING↗

NuGraph2 with explainability: post-hoc explanations for geometric neural network predictions

With the growing popularity of artificial intelligence (AI) used for scientific applications, the ability of attribute a result to a reasoning process from the network is in high demand for robust scientific generalizations to hold. In this work we aim to motivate the need for and demonstrate the use of post-hoc explainability methods when applied to AI methods used in scientific applications. To this end, we introduce explainability add-ons to the existing graph neural network (GNN) for neutrino tagging, NuGraph2. The explanations take the form of a suite of techniques examining the output of the network (node classifications) and the edge connections between them, and probing of the latent space using novel general-purpose tools applied to this network. We show how none of these methods are singularly sufficient to show network ‘understanding’, but together can give insights into the processes used in classification. While these methods are tested on the NuGraph2 application, they can be applied to a broad range of networks, not limited to GNNs. The code for this work is publicly available on GitHub at https://github.com/voetberg/XNuGraph.

Voetberg, Margaret [Fermilab] (ORCID:0009000527154↗

Novel symmetry-preserving neural network model for phylogenetic inference

Abstract Motivation Scientists world-wide are putting together massive efforts to understand how the biodiversity that we see on Earth evolved from single-cell organisms at the origin of life and this diversification process is represented through the Tree of Life. Low sampling rates and high heterogeneity in the rate of evolution across sites and lineages produce a phenomenon denoted “long branch attraction” (LBA) in which long nonsister lineages are estimated to be sisters regardless of their true evolutionary relationship. LBA has been a pervasive problem in phylogenetic inference affecting different types of methodologies from distance-based to likelihood-based. Results Here, we present a novel neural network model that outperforms standard phylogenetic methods and other neural network implementations under LBA settings. Furthermore, unlike existing neural network models in phylogenetics, our model naturally accounts for the tree isomorphisms via permutation invariant functions which ultimately result in lower memory and allows the seamless extension to larger trees. Availability and implementation We implement our novel theory on an open-source publicly available GitHub repository: https://github.com/crsl4/nn-phylogenetics.

59 BASIC BIOLOGICAL SCIENCES↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

tSFM 1.0: tRNA Structure–Function Mapper

Structure-conditioned information statistics have proven useful to predict and visualize tRNA Class-Informative Features (CIFs) and their evolutionary divergences. Although permutation P-values can quantify the significance of CIF divergences between two taxa, their naive Monte Carlo approximation is slow and inaccurate. The Peaks-over-Threshold approach of Knijnenburg et al. (2009) promises improvements to both speed and accuracy of permutation P-values, but has no publicly available API. Here, we present tRNA Structure–Function Mapper (tSFM) v1.0, an open-source, multi-threaded application that efficiently computes, visualizes and assesses significance of single- and paired-site CIFs and their evolutionary divergences for any RNA, protein, gene or genomic element sequence family. Multiple estimators of permutation P-values for CIF evolutionary divergences are provided along with confidence intervals. tSFM is implemented in Python 3 with compiled C extensions and is freely available through GitHub (https://github.com/tlawrence3/tSFM) and PyPI.

59 BASIC BIOLOGICAL SCIENCES↗