Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Python software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Adiabatic quantum linear regression

Abstract A major challenge in machine learning is the computational expense of training these models. Model training can be viewed as a form of optimization used to fit a machine learning model to a set of data, which can take up significant amount of time on classical computers. Adiabatic quantum computers have been shown to excel at solving optimization problems, and therefore, we believe, present a promising alternative to improve machine learning training times. In this paper, we present an adiabatic quantum computing approach for training a linear regression model. In order to do this, we formulate the regression problem as a quadratic unconstrained binary optimization (QUBO) problem. We analyze our quantum approach theoretically, test it on the D-Wave adiabatic quantum computer and compare its performance to a classical approach that uses the Scikit-learn library in Python. Our analysis shows that the quantum approach attains up to $${2.8 \times }$$ 2.8 × speedup over the classical approach on larger datasets, and performs at par with the classical approach on the regression error metric. The quantum approach used the D-Wave 2000Q adiabatic quantum computer, whereas the classical approach used a desktop workstation with an 8-core Intel i9 processor. As such, the results obtained in this work must be interpreted within the context of the specific hardware and software implementations of these machines.

97 MATHEMATICS AND COMPUTING↗

NISQ Benchmarking

Test suite of quantum algorithms for Noisy Intermediate Scale Quantum (NISQ) computers. The test suite includes benchmark-style code for quantum volume circuits (QV), fairness sampling circuits, quantum telecloning circuits, and other NISQ benchmark style algorithms on small problems (i.e., up to 100 qubits), such as Variational Quantum Eigensolver (VQE), Hamiltonian Simulation, and Grover unstructured search example circuits. These benchmark-style applications are implemented in quantum software packages, mostly IBM's QISKIT, but may include vendor-specific frameworks, such as PyQuil (for Rigetti) or Q\# for Microsoft, or CirQ (for Google) as the test suite grows with the vendor sample. The test suite also includes numerical simulation code for Quantum Alternating Operator Ansatz (QAOA) algorithms, VQE, Hamiltonian Simulation and search examples. Numerical simulation code simulates quantum computers on classical computers, which is only possible for small problem instances; the implementation framework of choice is typically within Python, using the numpy/scipy libraries as well as extensions to the Julia language.

Pelofske, Elijah↗

Strym: A Python Package for Real-time CAN Data Logging, Analysis and Visualization to Work with USB-CAN Interface

In this report, we describe a data analysis tool developed for decoding and analyzing vehicle data obtained from a passenger vehicle’s onboard controller area network (CAN) bus. The tool developed in this paper provides a timeseries framework to perform domain-specific analysis at scale when interpreting data from a vehicle or a collection of vehicles in light of how to design intelligent vehicle applications. The tool, called Strym, exploits the CAN bus mechanism of modern vehicles to capture data using commercially available CAN-to-USB hardware Comma.ai Panda devices, managed through open-source software Libpanda. Strym permits the decoding of vendor-specific CAN messages in a vehicle-agnostic manner. Through this, a researcher can characterize data throughput, assess data quality, and perform analyses. Such analyses are useful in a number of research such as studying human driving behavior in mixed-autonomy, new driver models, rare-event detection, traffic flow estimation, and custom control of vehicles.

Performance evaluation, Smart cities, Intelligent ↗

eQuilibrator 3.0: a database solution for thermodynamic constant estimation

Abstract eQuilibrator (equilibrator.weizmann.ac.il) is a database of biochemical equilibrium constants and Gibbs free energies, originally designed as a web-based interface. While the website now counts around 1,000 distinct monthly users, its design could not accommodate larger compound databases and it lacked a scalable Application Programming Interface (API) for integration into other tools developed by the systems biology community. Here, we report on the recent updates to the database as well as the addition of a new Python-based interface to eQuilibrator that adds many new features such as a 100-fold larger compound database, the ability to add novel compounds, improvements in speed and memory use, and correction for Mg2+ ion concentrations. Moreover, the new interface can compute the covariance matrix of the uncertainty between estimates, for which we show the advantages and describe the application in metabolic modelling. We foresee that these improvements will make thermodynamic modelling more accessible and facilitate the integration of eQuilibrator into other software platforms.

59 BASIC BIOLOGICAL SCIENCES↗

Implementation of Plot File Testing in the DYNA3D/ParaDyn Software Quality Assurance Suite

Automated testing of DYNA3D/ParaDyn plot files was added to the DYNA3D/ParaDyn software quality assurance (SQA) test suite. The new capability extracts select data from the plot files generated during each verification run and compares it to the same baseline answers used to verify the problem. Deviations between baseline answers and plot file values are reported in the same manner as solution discrepancies, and differences in precision levels between the baseline answers and plot file results are accounted for. The new testing leverages the existing SQA test suite framework and test problems and the Python Mili reader and minimally increases the overall run time (< 5%) of the SQA test suite. This new capability provides incremental end-toend testing of the most common DYNA3D/ParaDyn simulation workflows.

42 ENGINEERING↗

A generalized wind turbine cross section as a reduced-order model to gain insights in blade aeroelastic challenges

In this work, we present an approach to study the aeroelastic stability of a wind turbine by focusing on the dynamics of a blade cross section. We present a methodology to obtain a reduced-order model of the blade dynamics in the form of generalized cross-sectional quantities that approximates the aerodynamic and structural properties of the full blade. The motivation for the work is to gain a physical understanding of the influence of aerodynamic models such as dynamic wake and dynamic stall on the frequency and damping of the structure using a reduced-order model with low computational cost. The model may be coupled to two-dimensional computational fluid dynamics softwares or engineering unsteady airfoil aerodynamics models accounting for dynamic wake and dynamic stall. In the latter case, we can obtain monolithic state-space forms of the aeroelastic system of equations, which simplifies the determination of the modal parameters and therefore the study of stability. The work investigates wind turbines in operation or at standstill, where vortex-induced vibrations and stall-induced vibrations, respectively, might be an issue. The implementation is made available as part of the open-source Python package WELIB and as part of the open-source unsteady aerodynamic driver of OpenFAST.

17 WIND ENERGY↗

`SkyPy`: A package for modelling the Universe

SkyPy is an open-source Python package for simulating the astrophysical sky. It comprises a library of physical and empirical models across a range of observables and a command-line script to run end-to-end simulations. The library provides functions that sample realisations of sources and their associated properties from probability distributions. Simulation pipelines are constructed from these models using a YAML-based configuration syntax, while task scheduling and data dependencies are handled internally and the modular design allows users to interface with external software. SkyPy is developed and maintained by a diverse community of domain experts with a focus on software sustainability and interoperability. By fostering development, it provides a framework for correlated simulations of a range of cosmological probes including galaxy populations, large scale structure, the cosmic microwave background, supernovae and gravitational waves. Version 0.4 implements functions that model various properties of galaxies including luminosity functions, redshift distributions and optical photometry from spectral energy distribution templates. Future releases will provide additional modules, for example, to simulate populations of dark matter halos and model the galaxy-halo connection, making use of existing software packages from the astrophysics community where appropriate.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ThinCurr: An open-source 3D thin-wall eddy current modeling code for the analysis of large-scale systems of conducting structures

In this paper we present a new thin-wall eddy current modeling code, ThinCurr, for studying inductively-coupled currents in 3D conducting structures -- with primary application focused on the interaction between currents flowing in coils, plasma, and conducting structures of magnetically-confined plasma devices. The code utilizes a boundary finite element method on an unstructured, triangular grid to accurately capture device structures. The new code, part of the broader Open FUSION Toolkit, is open-source and designed for ease of use without sacrificing capability and speed through a combination of Python, Fortran, and C/C++ components. Scalability to large models is enabled through use of hierarchical off-diagonal low-rank compression of the inductance matrix, which is otherwise dense. Ease of handling large models of complicated geometry is further supported by automatic determination of supplemental elements through a greedy homology approach. Here, a detailed description of the numerical methods of the code and verification of the implementation of those methods using cross-code comparisons against the VALEN code and Ansys commercial analysis software is shown.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Integrating ORNL’s HPC and Neutron Facilities with a Performance-Portable CPU/GPU Ecosystem

We explore the development of a performance-portable CPU/GPU ecosystem to integrate two of the US Department of Energy’s (DOE’s) largest scientific instruments, the Oak Ridge Leadership Computing facility and the Spallation Neutron Source (SNS), both of which are housed at Oak Ridge National Laboratory. We select a relevant data reduction workflow use-case to obtain the differential scattering cross-section from data collected by SNS’s CORELLI and TOPAZ instruments. We compare the current CPU-only production implementation using the Garnet Python multiprocess package based on the Mantid C++ framework against our proposed CPU/GPU implementation that uses the LLVM-based, just-in-time Julia scientific language and the JACC.jl performance-portable package. Two proxy apps were developed: (i) an app for extracting relevant Mantid kernels (MDNorm) in C++ and (ii) the Julia MiniVATES.jl miniapp. We present performance results for NVIDIA A100 and AMD MI100 GPUs and AMD EPYC 7513 and 7662 CPUs. The results provide insights for future generations of data reduction software that can embrace performance portability for an integrated research infrastructure across DOE’s experimental and computational facilities.

Hahn, Steven↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

TPSAS-NF1676L-32060-DND

All previous PSP testing done in the Unitary Plan Wind Tunnel (UPWT) have required a significant amount of manual operation of the system. This has resulted in decreased testing efficiency and precluded the ability to provide near real-time data analysis to the customer. The overall goal of this project is to integrate the PSP data acquisition system into the supersonic UPWT data acquisition system (DAS) and create an adaptive software platform from which PSP data acquisition can be triggered by the tunnel and critical testing conditions can be recorded in real time for rapid analysis of the PSP data. This analysis includes the mapping of up to eight camera views onto a surface grid for analysis and converting to pressure using parameters supplied by the DAS. To fully implement this solution, communication must first be established between the Unitary DAS and the PSP DAS. This will be done by employing multiple scripts written in Python and C++ and implemented on a Linux cluster. These will be demonstrated and refined on an upcoming test (December 2018), and the successful completion will result in the ability to have automatic collection of PSP images and near real-time analysis capabilities.

Juliette Eddins↗

POST Explorer: A Design Space Exploration Tool for POST2

Recent improvements for the Program to Optimize Simulated Trajectories II (POST2) have included the development of an application programming interface (API). This API allows POST2 simulation inputs to be directly manipulated from other applications (such as MATLAB or Python), and the outputs from POST2 are streamed directly to the external application that enables visualization, data manipulation, etc. Through this framework, a new tool called POST Explorer is being developed that provides a user the capability to modify the simulation inputs and interrogate the outputs within the same application, with raw data inspection and visualization embedded. This tool can be leveraged for multiple types of analyses, such as parametric sweeps and sensitivity studies, and will be available with a future release of the POST2 software.

Robert Anthony Williams↗

POST Explorer: A Design Space Exploration Tool for POST2

Recent improvements for the Program to Optimize Simulated Trajectories II (POST2) have included the development of an application programming interface (API). This API allows POST2 simulation inputs to be directly manipulated from other applications (such as MATLAB or Python), and the outputs from POST2 are streamed directly to the external application that enables visualization, data manipulation, etc. Through this framework, a new tool called POST Explorer is being developed that provides a user the capability to modify the simulation inputs and interrogate the outputs within the same application, with raw data inspection and visualization embedded. This tool can be leveraged for multiple types of analyses, such as parametric sweeps and sensitivity studies, and will be available with a future release of the POST2 software.

Anthony Williams↗

Multi‐Decadal Decarbonization Pathways for U.S. Freight Rail

A‐STEP is a first‐of‐its‐kind, integrated, open‐source software tool aimed at guiding freight rail decarbonization decision‐making. It has tools for studying energy use details for individual trains, networks of trains, battery and hydrogen charging stations, national energy sourcing and pricing, and overall decarbonization costs and environmental impacts. It gives analysts an ability to study the challenges of making such change happen. Completely amenable to analyst specified inputs and parameter values, it can be customized to provide outputs for a wide variety of assumptions about future energy conditions and technological advances. Written in Python, C++, and VB.Net, A‐STEP can be implemented on both Windows and Linux‐based platforms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

AEcroscopy: A Software–Hardware Framework Empowering Microscopy Toward Automated and Autonomous Experimentation

Microscopy has been pivotal in improving the understanding of structure-function relationships at the nanoscale and is by now ubiquitous in most characterization labs. However, traditional microscopy operations are still limited largely by a human-centric click-and-go paradigm utilizing vendor-provided software, which limits the scope, utility, efficiency, effectiveness, and at times reproducibility of microscopy experiments. Here, in this work, a coupled software–hardware platform is developed that consists of a software package termed AEcroscopy (short for Automated Experiments in Microscopy), along with a field-programmable-gate-array device with LabView-built customized acquisition scripts, which overcome these limitations and provide the necessary abstractions toward full automation of microscopy platforms. The platform works across multiple vendor devices on scanning probe microscopes and electron microscopes. It enables customized scan trajectories, processing functions that can be triggered locally or remotely on processing servers, user-defined excitation waveforms, standardization of data models, and completely seamless operation through simple Python commands to enable a plethora of microscopy experiments to be performed in a reproducible, automated manner. This platform can be readily coupled with existing machine-learning libraries and simulations, to provide automated decision-making and active theory-experiment optimization to turn microscopes from characterization tools to instruments capable of autonomous model refinement and physics discovery.

47 OTHER INSTRUMENTATION↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

OpenSAMPL: An Open Source Library for Timing and Synchronization Measurements and Analytics

Today's power grid operators are implementing timing and synchronization solutions that provide resilience to Global Navigation Satellite System (GNSS) vulnerabilities. These vendor-specific solutions often come with additional software applications that are designed to monitor that vendor's synchronization performance data. However, resilient timing architectures often resulting in multi-vendor solutions, including approaches that blend terrestrial clocks with space-based subscription services. In such an environment, collecting, analyzing, and visualizing data from a variety of sources within a single platform was heretofore not possible. To address this need, the US Department of Energy's Center for Alternative Synchronization and Timing (CAST) developed OpenSAMPL, the Open Synchronized Analytics and Monitoring Platform, an open-source Python framework for processing, loading, and observing clock measurement data from distributed devices. OpenSAMPL enables the ingestion of diverse clock-probe sources into a scalable time-series database and applies robust analytics. OpenSAMPL currently supports two vendor data pipelines, and will be extended to more in the near future, enabling seamless monitoring of a variety of timing and synchronization devices in a common environment.

Grant, Josh [ORNL] (ORCID:0000000163475060)↗

Hindsight logging for model training

In modern Machine Learning, model training is an iterative, experimental process that can consume enormous computation resources and developer time. To aid in that process, experienced model developers log and visualize program variables during training runs. Exhaustive logging of all variables is infeasible, so developers are left to choose between slowing down training via extensive conservative logging, or letting training run fast via minimalist optimistic logging that may omit key information. As a compromise, optimistic logging can be accompanied by program checkpoints; this allows developers to add log statements post-hoc, and "replay" desired log statements from checkpoint---a process we refer to as hindsight logging. Unfortunately, hindsight logging raises tricky problems in data management and software engineering. Done poorly, hindsight logging can waste resources and generate technical debt embodied in multiple variants of training code. In this paper, we present methodologies for efficient and effective logging practices for model training, with a focus on techniques for hindsight logging. Our goal is for experienced model developers to learn and adopt these practices. To make this easier, we provide an open-source suite of tools for Fast Low-Overhead Recovery (flor) that embodies our design across three tasks: (i) efficient background logging in Python, (ii) adaptive periodic checkpointing, and (iii) an instrumentation library that codifies hindsight logging for efficient and automatic record-replay of model-training. Model developers can use each flor tool separately as they see fit, or they can use flor in hands-free mode, entrusting it to instrument their code end-to-end for efficient record-replay. Our solutions leverage techniques from physiological transaction logs and recovery in database systems. Evaluations on modern ML benchmarks demonstrate that flor can produce fast checkpointing with small user-specifiable overheads (e.g. 7%), and still provide hindsight log replay times orders of magnitude faster than restarting training from scratch.

Computer Science↗