Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Python codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

BatAnalysis - A Comprehensive Python Pipeline for Swift BAT Survey Analysis

The Swift Burst Alert Telescope (BAT) is a coded-aperture gamma-ray instrument with a large field of view that primarily operates in survey mode when it is not triggering on transient events. The survey data consist of 80- channel detector plane histograms that accumulate photon counts over periods of at least 5 minutes. These histograms are processed on the ground and are used to produce the survey data set between 14 and 195 keV. Survey data comprise >90% of all BAT data by volume and allow for the tracking of long-term light curves and spectral properties of cataloged and uncataloged hard X-ray sources. Until now, the survey data set has not been used to its full potential due to the complexity associated with its analysis and the lack of easily usable pipelines. Here, we introduce the BatAnalysis Python package, a wrapper for HEASoftpy, which provides a modern, opensource pipeline to process and analyze BAT survey data. BatAnalysis allows members of the community to use BAT survey data in more advanced analyses of astrophysical sources, including pulsars, pulsar wind nebula, active galactic nuclei, and other known/unknown transient events that may be detected in the hard X-ray band. We outline the steps taken by the Python code and exemplify its usefulness and accuracy by analyzing survey data of the Crab Nebula, NGC 2992, and a previously uncataloged MAXI transient. The BatAnalysis package allows for ~18 yr of BAT survey data to be used in a systematic way to study a large variety of astrophysical sources.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

BatAnalysis - A Comprehensive Python Pipeline for Swift BAT Survey Analysis

The Swift Burst Alert Telescope (BAT) is a coded-aperture gamma-ray instrument with a large field of view that primarily operates in survey mode when it is not triggering on transient events. The survey data consist of 80-channel detector plane histograms that accumulate photon counts over periods of at least 5 minutes. These histograms are processed on the ground and are used to produce the survey data set between 14 and 195 keV. Survey data comprise >90% of all BAT data by volume and allow for the tracking of long-term light curves and spectral properties of cataloged and uncataloged hard X-ray sources. Until now, the survey data set has not been used to its full potential due to the complexity associated with its analysis and the lack of easily usable pipelines. Here, we introduce the BatAnalysis Python package, a wrapper for HEASoftpy, which provides a modern, open-source pipeline to process and analyze BAT survey data. BatAnalysis allows members of the community to use BAT survey data in more advanced analyses of astrophysical sources, including pulsars, pulsar wind nebula, active galactic nuclei, and other known/unknown transient events that may be detected in the hard X-ray band. We outline the steps taken by the Python code and exemplify its usefulness and accuracy by analyzing survey data of the Crab Nebula, NGC 2992, and a previously uncataloged MAXI transient. The BatAnalysis package allows for ~18 yr of BAT survey data to be used in a systematic way to study a large variety of astrophysical sources.

79 ASTRONOMY AND ASTROPHYSICS↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Simulations of coherent scattering experiments at storage ring synchrotron radiation sources in the hard x-ray range

Detailed simulations of experiments carried out at modern light sources are directly related to the most efficient and productive use of these facilities for research in multiple branches of science and technology. The “Synchrotron Radiation Workshop” computer code with its Python interface, and Sirepo web-browser-based graphical user interface, currently supports physical optics simulations of coherent X-ray scattering and imaging experiments on user-defined virtual samples. We present examples of simulations of coherent scattering experiments that are typically performed at the Coherent Hard X-ray beamline at Brookhaven National Laboratory’s (BNL) National Synchrotron Light Source II. We also present several comparisons of the simulations with the results of actual coherent X-ray scattering experiments with nano-fabricated test samples produced at BNL’s Center for Functional Nanomaterials.

36 MATERIALS SCIENCE↗

PyMTensor 1.0.0

SAND2021-1654 O The Python Material Tensor code (PyMTensor) applies crystal symmetries to arbitrary material tensors. It uses symbolic algebra to exactly determine the zero tensor components and the relationships between the nonzero components for all crystallographic point groups. It is capable of computing components of higher-order material tensors needed when treating nonlinear effects. It has a simple and flexible interface for inputting tensor rank information and any known symmetries between tensor indices.Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Jensen, DanielS↗

Nvd Search To Stix

This code is a python based application that queries the National Vulnerability Database (NVD) API search term and CPE endpoints. It then sifts through the API response and uses the STIX2 python package to create STIX SDOs, SROs, and SCOs from the applicable data. If there are CWEs associated with the bundle, it queries the OpenCVE API for information on the weakness, then translates that data to STIX as well. It then combines all the data into a STIX bundle and outputs it to a JSON file.

Beckman, BryanR [Idaho National Laboratory (INL), ↗

Binder-benchmarking

SAND2025-07593O Binder-benchmarking evaluates the speed and memory impacts of C++, Python, and Matlab code binders. As a repository, it provides a way to locally run computation-based and memory-based benchmark suites on pybind11 and nanobind-based code in a Docker image. The software runs simple-speed and memory benchmarks on primitive navigation and integration exemplar algorithms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Walker II, Michael [Sandia National Lab. (SNL-CA),↗

SMART Deliverable 6.1.2a: Application of the ORION tool to the IBDP Carbon Storage Site

Forecasting and managing potential induced seismic activity is one of the challenges facing commercialscale geologic carbon sequestration (GCS), as well as other geologic energy extraction and byproduct disposal technologies. Historically, the process to develop robust, science-based forecasts of induced seismicity has required an integrated effort from experts in seismology, geomechanics, and reservoir engineering to manage data, develop and evaluate models of subsurface processes, and to calibrate and interpret the results from a range of models to understand site behavior relative to prescribed standards and in the context of uncertainty in geologic characterization data, forecasting models, and operational scenario uncertainty. The Operational Forecasting of Induced Seismicity (ORION) toolkit is an open-source, observation-based forecasting toolkit that is being co-developed by two U.S. DOE-funded initiatives: the National Risk Assessment Partnership (NRAP) and the Scienceinformed Machine Learning for Accelerating Real Time Decisions in Subsurface Applications (SMART) Initiative. ORION is designed to provide functionality to support decision making about seismic hazard analysis and risk management for GCS stakeholders ranging from the public to site operators to expert seismologists. The tool, which is written as open-source code in the Python programming language, is composed of a desktop graphical user interface (GUI) and an underlying forecasting engine. The forecasting engine uses available reservoir properties, well and fluid injection scenario details, and observed seismic catalog data as inputs to produce a set of temporal and spatio-temporal seismic forecasts.

58 GEOSCIENCES↗

Computing the Instantaneous Collision Probability between Satellites using Characteristic Function Inversion

The probability that two satellites overlap in space at a specified instant of time is called their instantaneous collision probability. Assuming Gaussian uncertainties and spherical satellites, this probability is the integral of a Gaussian distribution over a sphere. This paper shows how to compute the probability using an established numerical procedure called characteristic function inversion. The collision probability in the short-term encounter scenario is also evaluated with this approach, where the instant at which the probability is computed is the time of closest approach between the objects. Python and R code is provided to evaluate the probability in practice. Overall, the approach has been established for over fifty years, is implemented in existing software, does not rely on analytical approximations, and can be used to evaluate two and three dimensional collision probabilities.

79 ASTRONOMY AND ASTROPHYSICS↗

classLog: Logistic regression for the classification of genetic sequences

Introduction Sequencing and phylogenetic classification have become a common task in human and animal diagnostic laboratories. It is routine to sequence pathogens to identify genetic variations of diagnostic significance and to use these data in realtime genomic contact tracing and surveillance. Under this paradigm, unprecedented volumes of data are generated that require rapid analysis to provide meaningful inference. Methods We present a machine learning logistic regression pipeline that can assign classifications to genetic sequence data. The pipeline implements an intuitive and customizable approach to developing a trained prediction model that runs in linear time complexity, generating accurate output rapidly, even with incomplete data. Our approach was benchmarked against porcine respiratory and reproductive syndrome virus (PRRSv) and swine H1 influenza A virus (IAV) datasets. Trained classifiers were tested against sequences and simulated datasets that artificially degraded sequence quality at 0, 10, 20, 30, and 40%. Results When applied to a poor-quality sequence data, the classifier achieved between >85% to 95% accuracy for the PRRSv and the swine H1 IAV HA dataset and this increased to near perfect accuracy when using the full dataset. The model also identifies amino acid positions used to determine genetic clade identity through a feature selection ranking within the model. These positions can be mapped onto a maximum-likelihood phylogenetic tree, allowing for the inference of clade defining mutations. Discussion Our approach is implemented as a python package with code available at https://github.com/flu-crew/classLog .

Zeller, Michael A.↗

Heat Pump Retrofits for Central Plant Hydronic Heating Systems: A Software Toolkit for Screening and Design

Retrofitting existing central plants with high-efficiency heat pump technologies can play a crucial role in achieving long-term planning goals. Modern heat pump technologies are able to use waste heat recovery to meet a building's heating demand, but there is a lack of accessible tools designed for non-HVAC experts, such as building owners, to quickly and easily conduct what-if analysis, e.g., estimating retrofit costs and payback period for their partial or full equipment replacement. This paper introduces an open-source software toolkit designed to facilitate the initial screening and decision-making of heat pump retrofits in existing central plants using a building's yearly load profile from metered or utility bill data. The toolkit evaluates the technical and economic viability of replacing traditional central plant equipment with various options including water-to-water or air-to-water heat pumps, which can provide efficient and lower-cost heating and cooling. It allows users to compare current central plant configurations with retrofit scenarios, assessing energy consumption, life-cycle costs, and environmental impact. The toolkit offers (1) a web-based tool designed for user-friendly access by a broad audience and (2) Python-based source code for researchers and engineers conducting parametric studies and design parameter optimization. The toolkit compares a typical central plant configuration to a configuration that uses a heat pump to supply hydronic heating and cooling. The output metrics include energy consumption and output of each equipment, life-cycle cost analyses and metrics, and environmental impact of the system.

Excell, L↗

The InSAR Scientific Computing Environment

We have developed a flexible and extensible Interferometric SAR (InSAR) Scientific Computing Environment (ISCE) for geodetic image processing. ISCE was designed from the ground up as a geophysics community tool for generating stacks of interferograms that lend themselves to various forms of time-series analysis, with attention paid to accuracy, extensibility, and modularity. The framework is python-based, with code elements rigorously componentized by separating input/output operations from the processing engines. This allows greater flexibility and extensibility in the data models, and creates algorithmic code that is less susceptible to unnecessary modification when new data types and sensors are available. In addition, the components support provenance and checkpointing to facilitate reprocessing and algorithm exploration. The algorithms, based on legacy processing codes, have been adapted to assume a common reference track approach for all images acquired from nearby orbits, simplifying and systematizing the geometry for time-series analysis. The framework is designed to easily allow user contributions, and is distributed for free use by researchers. ISCE can process data from the ALOS, ERS, EnviSAT, Cosmo-SkyMed, RadarSAT-1, RadarSAT-2, and TerraSAR-X platforms, starting from Level-0 or Level 1 as provided from the data source, and going as far as Level 3 geocoded deformation products. With its flexible design, it can be extended with raw/meta data parsers to enable it to work with radar data from other platforms

geodetic imaging↗

An Automated Detection Methodology for Dry Well-Mixed Layers

The intense surface heating over arid land surfaces produces dry well-mixed layers (WML) via dry convection. These layers are characterized by nearly constant potential temperature and low, nearly constant water vapor mixing ratio. To further the study of dry WMLs, we created a detection methodology and supporting software to automate the identification and characterization of dry WMLs from multiple data sources including rawinsondes, remote sensing platforms, and model products. The software is a modular code written in Python, an open source language. Radiosondes from a network of synoptic stations in North Africa were used to develop and test the WML detection process. The detection involves an iterative decision tree that ingests a vertical profile from an input data file, performs a quality check for sufficient data density, and then searches upward through the column for successive points where the simultaneous changes in water vapor mixing ratio and potential temperature are less than the specified maxima. If points in the vertical profile meet the dry WML identification criteria, statistics are generated detailing the characteristics of each layer in the profile. At the end of the vertical profile analysis, there is an option to plot analyzed profiles in a variety of file formats. Initial results show that the detection methodology can be successfully applied across a wide variety of input data and North African environments and for all seasons. It is sensitive enough to identify dry WMLs from other types of isentropic phenomena such as subsidence layers and distinguish the current day’s dry WML from previous days.

Stephen D. Nicholls↗

Trajectory Simulation Using Multi Model Monte Carlo with Python (MXMCPy)

EDL (Entry, Descent and Landing) is the process from a vehicle approaching a surface to landing on it, such as a Mars rover approaching the planet before landing. POST2 (Program to Optimize Simulated Trajectories 2) is Langley’s primary EDL simulation tool and is used NASA-wide for simulations. POST2 can generate highly accurate results by running a precise, but time consuming, Monte Carlo (MC) simulation hundreds or thousands of times. Though POST2 can produce highly accurate results, it can take unrealistic time spans to generate these results, which has created a need to speed up the simulations. The new NASA software MXMCPy offers various ways to speed up the simulations while getting just as precise results. Instead of running high-precision POST2 simulations many times for traditional MC, MXMCPy can run fewer high-precision POST2 simulations and many less precise POST2 simulations and merge the results. MXMCPy contains 30+ different methods which will each suggest different allocations between model precision levels, which result in results of varying precision based on the POST2 simulation. I created Python and Bash code to automate the 5 steps of MXMCPy’s application to POST2. I also tested the precision of traditional Monte Carlo simulations to MXMCPy aided simulations and found that MXMCPy can achieve substantially more precise solutions at the same computer runtime. I learned Test Driven Development (TDD), a software programming workflow which involves writing computer-automated tests before writing the code which is being tested. These tests are ran every time the code is changed and they can find glitches in the code much quicker than a human can. This programming workflow saved me a lot of time because the automated tests could tell me exactly where the code had stopped working. I plan on using this software development method for future academic and professional software projects. I have greatly enjoyed my work at NASA, so I have been applying to NASA internships and Pathways positions. In addition, I plan on applying what I have learned about Test Driven Development to my computer science courses next semester

James Warner↗

Starshade Rendezvous: Exoplanet Sensitivity and Observing Strategy

Launching a starshade to rendezvous with the Nancy Grace Roman Space Telescope (Roman) would provide the first opportunity to directly image the habitable zones (HZs) of nearby sunlike stars in the coming decade. A report on the science and feasibility of such a mission was recently submitted to NASA as a probe study concept. The driving objective of the concept is to determine whether Earth-like exoplanets exist in the HZs of the nearest sunlike stars and have biosignature gases in their atmospheres. With the sensitivity provided by this telescope, it is possible to measure the brightness of zodiacal dust disks around the nearest sunlike stars and establish how their population compares with our own. In addition, known gas-giant exoplanets can be targeted to measure their atmospheric metallicity and thereby determine if the correlation with planet mass follows the trend observed in the Solar System and hinted at by exoplanet transit spectroscopy data. We provide the details of the calculations used to estimate the sensitivity of Roman with a starshade and describe the publicly available Python-based source code used to make these calculations. Given the fixed capability of Roman and the constrained observing windows inherent for the starshade, we calculate the sensitivity of the combined observatory to detect these three types of targets, and we present an overall observing strategy that enables us to achieve these objectives.

Andrew Frederic Romero-wolf↗