Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Python codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Geothermal Play Fairway Analysis for Low-Temperature Resources in the Denver Basin

This dataset is part of an effort to highlight the advantages of incorporating low-temperature (< 150 C) geothermal resource evaluation into the implementation of combined heat and power (CHP), and geothermal direct use (GDU) technologies (e.g., space heating and/or cooling). For this Denver Basin example, resource favorability maps were created to identify potentially favorable areas for further geothermal exploration and are provided here. Favorability was based on three types of data: (1) geologic, (2) economic, and (3) risk. This raw data is also provided below. Geologic data include bottom-hole temperatures (BHT) from oil and gas wells, water co-production volumes from oil and gas wells, well groundwater levels, hot spring locations, temperatures, and chemistries, faults, and earthquakes. Economic feasibility data include population, thermal energy demand, infrastructure, and roads. Risk data (which includes data on excluded areas) include flood plains, protected lands (e.g. wildlife conservation areas, national parks). The included report describes this project in detail, covering workflows, relevant datasets, Python code, and both common and composite maps used to create low-temperature geothermal resource favorability maps for the Denver Basin, which extends across Colorado, Nebraska, and Wyoming. The figures in this report include: maps of the original datasets; maps of transformed data and derived parameters (such as the geothermal gradient or thermal conductivity); results of uncertainty analyses; results of data completeness (using the GeoRePORT tool); examples of the data combination and processing (using the geoPFA Python library, which is introduced in the attached report); favorability maps for each criteria; and a final combined favorability map. This project is designed to facilitate future deployment of CHP and GDU by providing data, tools, and a workflow applicable to low-temperature geothermal resources in sedimentary basins.

15 GEOTHERMAL ENERGY↗

Review of Grey Box/Black Box Data Contamination Metrics on Open and Commercial Models

Dataset contamination is a problem where benchmarks and tasks used to evaluate the capabilities of Large Language Models (LLMs) have been incorporated into the training dataset of the models. This gives a false sense of performance that can overestimate how these models will function on truly unseen data. This problem becomes worse with commercial LLMs with larger and non-accessible training data, so techniques have been developed to try to measure the degree to which a model is contaminated with a benchmark’s data. To understand the effectiveness of these techniques, particularly when evaluating contamination on coding tasks, we review trends and categorize techniques by the degree of access to the model that is required. The research literature on this topic has reported mixed effectiveness of these techniques, so we select a set of black box (text access only) and grey box (access to model loss/probabilities required) techniques and apply them to both commercial and non-commercial models. We implement these metrics as part of a framework to test the contamination of Python code in LLMs to see to what extent we can replicate the effectiveness (or ineffectiveness) of these contamination detection techniques. Though we find mixed results in the capabilities of these metrics to identify contamination, we do observe evidence that they can identify contamination (broadly) in fine-tuned models when both a baseline and fine-tuned model is present. Additionally, similarity metrics were able to identify between contaminated and uncontaminated data even in situations where the data is distributionally similar (e.g., drawn from the same set of code projects).

97 MATHEMATICS AND COMPUTING↗

TUMME: Tsinghua University Minnesota Master Equation program

We report that TUMME is a program for assembling and solving master equations for gas-phase chemical kinetics based on chemically significant eigenmodes. TUMME has interfaces to the Gaussian, Polyrate, and/or MSTor output files that allow the master equation code to obtain the microcanonical flux coefficients needed for the coefficient matrix of the master equation. The flux coefficients for reactions with barriers can be calculated by multi-structural variational transition state theory with small-curvature tunneling (MS-VTST/SCT) or by simpler approximations to this such as conventional transition state theory without tunneling (also called RRKM theory). The flux coefficients for barrierless reactions are provided by a hard-sphere model. TUMME is written in double precision with Python 3; quadruple and octuple precision are also available for some subtasks in C++. The Python code can run in serial or parallel (MP or MPI), and the C++ code can run on a single processor or on multiple processors with OpenMP.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evidence-based Graph Adversary Mapping (EGRAM) [Poster]

Cybersecurity companies such as CrowdStrike, Dragos, Microsoft and Unit 42 categorize Advanced Persistent Threats (APTs) using their own naming schemes. As a result, these APTs are mapped to different malware sources and campaigns, all from differing sources, leading to inconsistent mapping. Inconsistent mapping causes confusion and adds further obscurity around these groups, making it difficult to track and mitigate APT cyberattacks. The Evidence-based Graph Adversary Mapping (EGRAM) tool remediates the mapping challenge by collecting, updating and converting adversary data and their sources into a valid, codified STIX v2.1 bundle which is then stored in a Neo4j graph database. It utilizes graph traversal methods and centrality analysis to generate actionable information as a Structured Threat Intelligence Graph (STIG), based on user queries. EGRAM exists as Python code and a Jupyter Notebook that acts as a searchable, evidence-based, source of intelligence for APT groups’ artifacts and cyber campaigns.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

ORT: a workflow linking genome-scale metabolic models with reactive transport codes

Abstract Motivation Nutrient and contaminant behavior in the subsurface are governed by multiple coupled hydrobiogeochemical processes which occur across different temporal and spatial scales. Accurate description of macroscopic system behavior requires accounting for the effects of microscopic and especially microbial processes. Microbial processes mediate precipitation and dissolution and change aqueous geochemistry, all of which impacts macroscopic system behavior. As ‘omics data describing microbial processes is increasingly affordable and available, novel methods for using this data quickly and effectively for improved ecosystem models are needed. Results We propose a workflow (‘Omics to Reactive Transport—ORT) for utilizing metagenomic and environmental data to describe the effect of microbiological processes in macroscopic reactive transport models. This workflow utilizes and couples two open-source software packages: KBase (a software platform for systems biology) and PFLOTRAN (a reactive transport modeling code). We describe the architecture of ORT and demonstrate an implementation using metagenomic and geochemical data from a river system. Our demonstration uses microbiological drivers of nitrification and denitrification to predict nitrogen cycling patterns which agree with those provided with generalized stoichiometries. While our example uses data from a single measurement, our workflow can be applied to spatiotemporal metagenomic datasets to allow for iterative coupling between KBase and PFLOTRAN. Availability and implementation Interactive models available at https://pflotranmodeling.paf.subsurfaceinsights.com/pflotran-simple-model/. Microbiological data available at NCBI via BioProject ID PRJNA576070. ORT Python code available at https://github.com/subsurfaceinsights/ort-kbase-to-pflotran. KBase narrative available at https://narrative.kbase.us/narrative/71260 or static narrative (no login required) at https://kbase.us/n/71260/258. Supplementary information Supplementary data are available at Bioinformatics online.

54 ENVIRONMENTAL SCIENCES↗

PyOMP: Multithreaded Parallel Programming in Python

We know that Python is a widely used language in scientific computing. When the goal is high performance, however, Python lags far behind low-level languages such as C and Fortran. To support applications that stress performance, Python needs to access the full capabilities of modern CPUs. That means support for parallel multithreading. In this paper, we describe PyOMP, a system that enables OpenMP in Python. Programmers write code in Python with OpenMP, Numba generates code that compiles to LLVM, and the resulting programs run with performance that approaches that from code written with C and OpenMP. In this paper we provide an update on the PyOMP project and explain how to install it and use it to write parallel multithreaded code in Python.

97 MATHEMATICS AND COMPUTING↗

An Open-Source Numerical Model for Mitigating Refractory Alloy Hot Cracking Susceptibility

Refractory alloys are susceptible to solidification cracking during welding and 3D printing. Composition control is an effective method of controlling solidification cracking. This work evaluates the effect of compositional variation in refractory metal systems on a computed solidification cracking susceptibility. A numerical model has been developed using Python code and open-source CALPHAD software to calculate Kou’s crack susceptibility index. The model is validated against past weldability studies performed on several refractory alloy systems. The approach is extended towards the development of new alloys with improved 3D printability and weldability and is shown to have utility in defining compositional limits for existing alloys and feedstocks. Furthermore, the model will aid in determining process controls for powder reuse and recycling.

pycalphad↗

Data on Cu- and Ni-Si-Mn-rich solute clustering in a neutron irradiated austenitic stainless steel

The data presented in this article is supplementary to the research article “Phase instabilities in austenitic steels during particle bombardment at high and low dose rates” (Levine et al.). Needle-shaped samples were prepared with focused ion beam milling from a 304L stainless steel that was irradiated with fast neutrons (E 0.1 MeV) in the BOR-60 reactor at 318 °C to 47.5 dpa. Atom probe tomography (APT) experiments in voltage mode were then conducted on a Cameca LEAP 5000X HR. Atom position, range, and mass spectrum files after reconstruction with Cameca’s IVAS software are included. Cu- and Ni-Si-Mn-rich solute nanoclusters were identified and analyzed using the Open Source Characterization of APT Reconstructions (OSCAR) program. Python code for OSCAR, information on the program’s underlying algorithm, and sample output files are provided. A proximity histogram of a Ni-Si-Mn-rich cluster and a 1D density/solute concentration profile of a Cu-rich cluster are given to demonstrate OSCAR’s analytical functionalities. The provided APT dataset is valuable for benchmarking phase instabilities in neutron-irradiated austenitic stainless steels that occur at high doses. The OSCAR program can be reused to process other APT data sets where solute nanoclustering is of interest.

42 ENGINEERING↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array↗

Combining Astrometry and Elemental Abundances: The Case of the Candidate Pre-Gaia Halo Moving Groups G03-37, G18-39, and G21-22

While most moving groups are young and nearby, a small number have been identified in the Galactic halo. Understanding the origin and evolution of these groups is an important piece of reconstructing the formation history of the halo. Here we report on our analysis of three putative halo moving groups: G03-37, G18-39, and G21-22. Based on Gaia EDR3 data, the stars associated with each group show some scatter in velocity (e.g., Toomre diagram) and integrals of motion (energy, angular momentum) spaces, counter to expectations of moving-group stars. We choose the best candidate of the three groups, G21-22, for follow-up chemical analysis based on high-resolution spectroscopy of six presumptive members. Using a new Python code that uses a Bayesian method to self-consistently propagate uncertainties from stellar atmosphere solutions in calculating individual abundances and spectral synthesis, we derive the abundances of α- (Mg, Si, Ca, Ti), Fe-peak (Cr, Sc, Mn, Fe, Ni), odd-Z (Na, Al, V), and neutron-capture (Ba, Eu) elements for each star. We find that the G21-22 stars are not chemically homogeneous. Based on the kinematic analysis for all three groups and the chemical analysis for G21-22, we conclude the three are not genuine moving groups. The case for G21-22 demonstrates the benefit of combining kinematic and chemical information in identifying conatal populations when either alone may be insufficient. Comparing the integrals of motion and velocities of the six G21-22 stars with those of known structures in the halo, we tentatively associate them with the Gaia-Enceladus accretion event.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Computational and Experimental Investigation of Chiral and Achiral Two‐Dimensional Organic Lead Bromide Perovskites: Octahedral Distortions and Electronic and Optical Properties

A computational investigation is presented, in conjunction with synthesis and experimental characterization, into the structural, electronic, and optical properties of layered two-dimensional organic lead bromide perovskites. Materials based on the chiral (R/S)-4-fluoro-α-methylbenzylammonium (R/S-FMBA), which have been shown to lead to bright room-temperature circularly polarized luminescence, are contrasted with the similar achiral 4-fluorobenzylammonium (FBA). Using density functional theory (DFT) with van der Waals (vdW) corrections, relaxed structures (compared with X-ray diffraction, XRD) and optical absorption spectra (compared with experiments) are studied, as well as band structure and orbital character of transitions. A Python code is developed and provided to calculate octahedral distortions and compare DFT and XRD results, finding that vdW corrections are important for accuracy and that DFT overestimates octahedral tilt angles. (FMBA) 2 PbBr 4 shows among the largest tilt angle differences (often termed Δ β ) reported, 14°–15°, indicating strong inversion symmetry-breaking, which enables its chiral emission. A large resulting Dresselhaus spin-splitting effect is found. The lowest-energy optical transitions involve the perovskite only and are polarized within the layer. This work furthers understanding of structure-property relations with applications to optoelectronics and spintronics.

UV/vis spectroscopy↗

Gap-filling eddy covariance methane fluxes: Comparison of machine learning model predictions and uncertainties at FLUXNET-CH4 wetlands

Time series of methane fluxes measured by eddy-covariance require gap-filling to estimate annual emissions. Gap-filling methane fluxes is challenging because of high variability and complex responses to multiple drivers. To date, there is no widely established gap-filling standard for methane, with regards both to the best model algorithms and predictors. In this study, we address the need for standardization by synthesizing results of gap-filling methods applied at 17 wetland sites spanning boreal to tropical regions including all major wetlands classes and two rice paddies. We introduce new procedures for: 1) creating realistic artificial gap scenarios, 2) training and evaluating gap-filling models without overstating performance, and 3) predicting half-hourly methane fluxes and annual emissions with robust uncertainty estimates. We tested a conventional method (marginal distribution sampling) and four machine learning algorithms - penalized linear regression, artificial neural networks, random forests, and boosted decision trees - and four predictor sets, including temporal, meteorological, ecosystem carbon and energy flux, and soil predictors. We find that the conventional method can achieve similar median performance to the machine learning models but is worse than the best machine learning models and relatively insensitive to predictor choices. Of the machine learning models, decision tree algorithms performed the best in cross-validation experiments, even with a baseline predictor set, and artificial neural networks showed comparable performance when using all predictors. Soil temperature was frequently the most important predictor whilst water table depth was important at sites with substantial water table fluctuations, highlighting the value of data on soil conditions. Raw gap-filling uncertainties from the machine learning models were underestimated and we propose a method to calibrate uncertainties to observations. Finally, we gap-fill and provide summary evaluation metrics for all 81 sites in the FLUXNET-CH4 community dataset and publicly release the python code for model development, evaluation, and uncertainty estimation.

42 ENGINEERING↗

Solution and sensitivity analysis of nonlinear equations using a hypercomplex-variable Newton-Raphson method

Here, the classical Newton-Raphson (NR) method for solving nonlinear equations is enhanced in two ways through the use of hypercomplex variables and algebra. In particular, i) the Jacobian is computed in a highly accurate and automated way, and ii) the derivative of the solution to the nonlinear equations is computed with respect to any parameter contained within the system of equations. These advances provide two significant enhancements in that it is straightforward to provide an accurate Jacobian and to construct a reduced order model (ROM) of arbitrary order with respect to any parameter of the system. The ROM can then be used to approximate the solution for other parameter values without requiring additional solutions of the nonlinear equations. Several case studies are presented including 1D and 2D academic examples with fully functioning Python code provided. Additionally, a case of study of the catenary of an elastic cable subject to its own weight and a vertical point load. Derivatives up to 10th order were computed with respect to material, loading, and geometrical parameters. The derivatives were used to generate reduced order models of the cable deformation and reaction forces at its ends with respect to multiple input parameters. Results show that from a single hypercomplex evaluation of the cable under a single vertical point load, it is possible to generate an accurate reduced order model capable of predicting the cable deformation with 1.5 times the load in the opposite direction and with 3.5 times the load in the same direction without resolving the system of equations.

97 MATHEMATICS AND COMPUTING↗

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Short-term electricity load forecasting: Application-driven evaluation of machine learning models across spatial and temporal scales

As we transition towards a decarbonized economy, the integration of variable renewable energy resources and new demands (e.g., electric vehicles, heat pumps) into the electricity grid places unprecedented pressure on grid operators to effectively anticipate and manage peak load. In this context, machine learning algorithms are proving to be indispensable for accurate short-term load forecasting, a crucial task to address these challenges. This study benchmarks 6 machine learning algorithms, including three neural networks and three tree-based algorithms, across various levels of spatial aggregation and time horizons (1, 4, 8, 24, and 48 h). The central contribution of this work is the comparison and analysis of load forecasting models not only based on statistical metrics, but also based on a novel error metric, which evaluates the cost implications of forecast errors for power system stakeholders. Results show that tree-based models outperform neural networks, based on statistical metrics, and yield less skewed error distributions for most spatial scales. However, through the lens of the novel error metric, neural networks are the more competitive choice, especially for forecast horizons that exceed 8 h. The study concludes with actionable recommendations to grid operators and highlights the need for the development of error metrics that link forecasting accuracy to operational costs. To promote transparency and open science, the datasets and Python code are open-sourced via a supplementary repository.

Houben, Nikolaus↗

Economic assessment of seismic monitoring for underground hydrogen storage

Underground hydrogen storage (UHS) plays a key role in the energy landscape. However, like other subsurface engineering technologies, UHS may cause leakage into the groundwater or atmosphere and possibly induce local seismicity. To reduce these risks, seismic monitoring could be a viable technique to track the UHS plume, detect leakages, and locate induced seismicity events. Seismic monitoring has been proposed to safely monitor UHS, but research in this area is still new and requires field studies. Lab and theoretical studies have demonstrated the validity of seismic monitoring for UHS. Therefore, it is imperative to analyze the economic feasibility of seismic monitoring for UHS. Hence, we develop a cost model and open-source Python code for seismic monitoring that considers types of seismometers, comprehensive operational scenarios, detection thresholds, and long-term leakage monitoring. A case study is further provided to validate the cost model on reservoir simulations of UHS. We find that the levelized cost for a 10-year operating UHS site will range on the order of ∼0.003 $\$$/kg. The methods developed in this study could also be applied to the monitoring of groundwater, gas, and/or wastewater injection.

08 HYDROGEN↗

GMFOLD: Subgraph matching for high-throughput DNA-aptamer secondary structure classification and machine learning interpretability

Aptamers are oligonucleotide receptors that bind to their targets with high affinity. Here, we consider aptamers comprised of single-stranded DNA that undergo target-binding-induced conformational changes, giving rise to unique secondary and tertiary structures. Given a specific aptamer primary sequence, there are well-established computational tools (notably mfold) to predict the secondary structure via free energy minimization algorithms. While mfold generates secondary structures for individual sequences, there is a need for a high-throughput process whereby thousands of DNA structures can be predicted in real-time for use in an interactive setting, when combined with aptamer selections that generate candidate pools that are too large to be experimentally interrogated. We developed a new Python code for high-throughput aptamer secondary structure determination (GMfold). GMfold uses subgraph matching methods to group aptamer candidates by secondary structure similarities. We also improve an open-source code, SeqFold, to incorporate subgraph matching concepts. We represent each secondary structure as a lowest-energy bipartite subgraph matching of the DNA graph to itself. These new tools enable thousands of DNA sequences to be compared based on their secondary structures, using machine-learning algorithms. This process is advantageous when analyzing sequences that arise from aptamer selections via systematic evolution of ligands by exponential enrichment (SELEX). This work is a building block for future machine-learning-informed DNA-aptamer selection processes to identify aptamers with improved target affinity and selectivity and advance aptamer biosensors and therapeutics.

Aptamer↗