Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mathematical software performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Adiabatic quantum linear regression

Abstract A major challenge in machine learning is the computational expense of training these models. Model training can be viewed as a form of optimization used to fit a machine learning model to a set of data, which can take up significant amount of time on classical computers. Adiabatic quantum computers have been shown to excel at solving optimization problems, and therefore, we believe, present a promising alternative to improve machine learning training times. In this paper, we present an adiabatic quantum computing approach for training a linear regression model. In order to do this, we formulate the regression problem as a quadratic unconstrained binary optimization (QUBO) problem. We analyze our quantum approach theoretically, test it on the D-Wave adiabatic quantum computer and compare its performance to a classical approach that uses the Scikit-learn library in Python. Our analysis shows that the quantum approach attains up to $${2.8 \times }$$ 2.8 × speedup over the classical approach on larger datasets, and performs at par with the classical approach on the regression error metric. The quantum approach used the D-Wave 2000Q adiabatic quantum computer, whereas the classical approach used a desktop workstation with an 8-core Intel i9 processor. As such, the results obtained in this work must be interpreted within the context of the specific hardware and software implementations of these machines.

97 MATHEMATICS AND COMPUTING↗

OpenCHAMI Developer Summit [Slides]

The mission of the OpenCHAMI consortium is to steward the collaborative development and continuous evolution of cloud-like software to manage High Performance Computing capacity regardless of the size or deployment platform. We are guided by the operators and practitioners who use modern tooling and concepts to address the needs of classical HPC applications and the growing AI/ML and Data Science community that wish to leverage HPC capacity within their own workflows, to meet their needs with their own tools.

97 MATHEMATICS AND COMPUTING↗

Early Application Results on Pre-exascale Architecture with Analysis of Performance Challenges and Projections (Milestone PM-AD-1080 WBS 2.2)

This Exascale Computing Project (ECP) Milestone Report summarizes the status of all 30 ECP Applications Development (AD) sub-projects at the end of FY19. In August and September of 2019, a comprehensive assessment of AD projects was conducted jointly by the ECP leadership and a team of external subject matter experts. Reviews took place in person over five days-two at the National Renewable Energy Laboratory and three at Argonne National Laboratory and the University of Chicago. The review committees were tasked with evaluating each sub-project's progress in porting their code(s) to current multi-GPU architectures considered precursors to planned exascale machines. This includes characterizing which modules have been ported to multi-accelerator nodes, initial performance analyses, the status of software integration, and a current vision of successes, obstacles, and next steps. As such this report contains not only an accurate snapshot of each sub-project's current status, but also represents an unprecedentedly broad account of experiences porting large scientific applications to next-generation HPC architectures.

97 MATHEMATICS AND COMPUTING↗

TAO Users Manual (Rev. 3.15)

The Toolkit for Advanced Optimization (TAO) focuses on the development of algorithms and software for the solution of large-scale optimization problems on high-performance architectures. Areas of interest include unconstrained and bound-constrained optimization, nonlinear least squares problems, optimization problems with partial differential equation constraints, and variational inequalities and complementarity constraints. The development of TAO was motivated by the scattered support for parallel computations and the lack of reuse of external toolkits in current optimization software. Our aim is to produce high-quality optimization software for computing environments ranging from workstations and laptops to massively parallel high-performance architectures. Our design decisions are strongly motivated by the challenges inherent in the use of large-scale distributed memory architectures and the reality of working with large, often poorly structured legacy codes for specific applications.

97 MATHEMATICS AND COMPUTING↗

Task Parallelism to Optimize Performance of Environmental Modeling Software

Climate modeling is an integral part of environmental research, from studying rare phenomena to predicting future climate trends. The need for more accurate models is only growing, but as climate modeling capabilities advance, existing workflows require optimization to recoup performance. A solution comes in the form of task parallelism, a novel programming capability that provides an opportunity for optimization at execution time by allowing tasks to be executed in parallel, reducing runtime significantly. Using Parsl, an intuitive and scalable parallel scripting library for Python, we implement task parallelism within support software to aid in the continuous advancement of climate modeling technology.

54 ENVIRONMENTAL SCIENCES↗

Jas4pp — A data-analysis framework for physics and detector studies

This paper describes the Jas4pp framework for exploring physics cases and for detector-performance studies of future particle collision experiments. Jas4pp is a multi-platform Java program for numeric calculations, scientific visualization in 2D and 3D, storing data in various file formats and displaying collision events and detector geometries. It also includes complex data-analysis algorithms for function minimization, regression analysis, event reconstruction (such as jet reconstruction), limit settings and other libraries widely used in particle physics. The framework can be used with several scripting languages, such as Python/Jython, Groovy and JShell. Several benchmark tests discussed in the paper illustrate significant improvements in the performance of the Groovy and JShell scripting languages compared to the standard Python implementation in C. Furthermore, the improvements for numeric computations in Java are attributed to recent enhancements in the Java Virtual Machine.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Map Applications to Target Exascale Architecture with Machine-Specific Performance Analysis, Including Challenges and Projections

This Exascale Computing Project (ECP) milestone report summarizes the status of all 30 ECP Applications Development (AD) subprojects at the end of FY20. In October and November of 2020, a comprehensive assessment of AD projects was conducted by the ECP leadership. Reviews occurred virtually between October 27, 2020 and November 12, 2020. The review committee—consisting of the AD lead, deputy, and L3—was tasked with evaluating each subproject’s progress in porting their codes to early exascale architectures considered precursors to the planned exascale machines. This includes characterizing which modules have been ported to multi-accelerator nodes, initial performance analyses, the status of software integration, and a current vision of successes, obstacles, and next steps. As such, this report contains not only an accurate snapshot of each subproject’s current status but also represents an unprecedentedly broad account of experiences in porting large scientific applications to next-generation high-performance computing architectures.

97 MATHEMATICS AND COMPUTING↗

Developing a Hybrid Electric Vehicle Eco-Cooperative Adaptive Cruise Control System at Signalized Intersections.

This study develops an eco-driving strategy for hybrid electric vehicles (HEVs) in the vicinity of signalized intersections, entitled HEV Eco-Cooperative Adaptive Cruise Control at Intersections (Eco-CACC-I). The proposed system computes real-time, energy-optimized vehicle trajectories using HEV vehicle dynamics and energy consumption models. In the proposed system, a simple HEV energy model is used to compute the instantaneous fuel consumption. This HEV energy model is selected since it is general, transferable, and can be easily used to compute instantaneous energy consumption levels for HEVs without the additional input of vehicle engine data or complicated power control strategies. In addition, a vehicle dynamics model is used to capture the relationship between speed, acceleration level, and tractive/resistance forces on vehicles. The energy-optimum problem is formulated as an optimization problem with constraints, which is solved using a moving-horizon dynamic programming approach. The proposed HEV Eco-CACC-I system was tested to evaluate its performance for various speed limits, roadway grades, and signal timings. Lastly, the proposed HEV controller was implemented in a microscopic traffic simulation software to test its network-wide performance. The test results from an arterial corridor with three signalized intersections demonstrate that the proposed system can effectively reduce stop-and-go traffic in the vicinity of signalized intersections producing savings of 7.4% in energy consumption, 5.8% in traffic delay and 23% vehicle stops, respectively.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Methods and Experiences for Developing Abstractions for Data-intensive, Scientific Applications

Developing software for scientific applications that require the integration of diverse types of computing, instruments, and data present challenges that are distinct from commercial software. These applications require scale, and the need to integrate various programming and computational models with evolving and heterogeneous infrastructure. Pervasive and effective abstractions for distributed infrastructures are thus critical; however, the process of developing abstractions for scientific applications and infrastructures is not well understood. While theory-based approaches for system development are suited for well-defined, closed environments, they have severe limitations for designing abstractions for scientific systems and applications. The design science research (DSR) method provides the basis for designing practical systems that can handle real-world complexities at all levels. In contrast to theory-centric approaches, DSR emphasizes both practical relevance and knowledge creation by building and rigorously evaluating all artifacts. In this work, we show how DSR provides a well-defined framework for developing abstractions and middleware systems for distributed systems. Specifically, we address the critical problem of distributed resource management on heterogeneous infrastructure over a dynamic range of scales, a challenge that currently limits many scientific applications. We use the pilot-abstraction, a widely used resource management abstraction for high-performance, high throughput, big data, and streaming applications, as a case study for evaluating the DSR activities. For this purpose, we analyze the research process and artifacts produced during the design and evaluation of the pilot-abstraction. We find DSR provides a concise framework for iteratively designing and evaluating systems. Finally, we capture our experiences and formulate different lessons learned.

97 MATHEMATICS AND COMPUTING↗

Open-Source Steady-State Models for Integration of Wave Energy Converter into Microgrids

This paper proposes a software framework, WEC-Grid, for integrating wave energy converters (WECs) into power flow software, such as Siemens PSS®E, to aid the integration of alternative energy sources into Microgrids. While integrating alternative sources such as WECs presents specific challenges such as cost, power quality, and power variability, wave energy is a promising renewable energy resource. Evaluating the integration of WECs into the power grid is a complex and nuanced problem that requires seamless communication between a WEC model and power flow software. The presented WEC-Grid software framework bridges and extends the functionality of WEC-Sim, an open-source WEC modeling package for MATLAB, through a wave-to-wire (W2W) electro-mechanical power conversion and processing model. WEC-Grid acts as a software wrapper, handler, and communication layer between the W2W modeler and power flow software. The software is designed to represent each grid system as a class object, allowing power system operators to perform power system duties such as contingency planning and dispatch operations. The integration of WECs with PSS®E’s power flow calculations workflow is demonstrated with an IEEE RTS case study.

16 TIDAL AND WAVE POWER↗

Numerical methods and hypoexponential approximations for gamma distributed delay differential equations

Abstract Gamma distributed delay differential equations (DDEs) arise naturally in many modelling applications. However, appropriate numerical methods for generic gamma distributed DDEs have not previously been implemented. Modellers have therefore resorted to approximating the gamma distribution with an Erlang distribution and using the linear chain technique to derive an equivalent system of ordinary differential equations (ODEs). In this work, we address the lack of appropriate numerical tools for gamma distributed DDEs in two ways. First, we develop a functional continuous Runge–Kutta (FCRK) method to numerically integrate the gamma distributed DDE without resorting to Erlang approximation. We prove the fourth-order convergence of the FCRK method and perform numerical tests to demonstrate the accuracy of the new numerical method. Nevertheless, FCRK methods for infinite delay DDEs are not widely available in existing scientific software packages. As an alternative approach to solving gamma distributed DDEs, we also derive a hypoexponential approximation of the gamma distributed DDE. This hypoexponential approach is a more accurate approximation of the true gamma distributed DDE than the common Erlang approximation but, like the Erlang approximation, can be formulated as a system of ODEs and solved numerically using standard ODE software. Using our FCRK method to provide reference solutions, we show that the common Erlang approximation may produce solutions that are qualitatively different from the underlying gamma distributed DDE. However, the proposed hypoexponential approximations do not have this limitation. Finally, we apply our hypoexponential approximations to perform statistical inference on synthetic epidemiological data to illustrate the utility of the hypoexponential approximation.

97 MATHEMATICS AND COMPUTING↗

Seascape Interface Control Document

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source software, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Automated Vulnerability Detection (AVUD) for Compiled Smart Grid Software

This project developed and implemented a system for conducting cybersecurity vulnerability detection of smart grid components and systems by performing static analysis of compiled software (“firmware”). The resulting system for automated vulnerability detection (AVUD) was implemented as part of Oak Ridge National Laboratory’s existing test bed for smart meters, the Sustainable Campus Initiative. The work consisted of two phases: the first phase implemented the necessary software and computational models to perform the analysis, and the second phase demonstrated the system on example firmware in partnership with smart meter manufacturer Sensus USA, Inc. The resulting system won an R&D 100 award and has been successfully commercialized, winning a National Laboratory Consortium Commercialization Award.

97 MATHEMATICS AND COMPUTING↗

Introducing PRIMRE's MRE Software Knowledge Hub (February 2021)

This paper focuses on the role of the Marine Renewable Energy (MRE) Software Knowledge Hub on the Portal and Repository for Information on Marine Renewable Energy (PRIMRE). The MRE Software Knowledge Hub provides online services for MRE software users and developers, and seeks to develop assessments and recommendations for improving MRE software in the future. Online software discovery platforms, known as the Code Hub and the Code Catalog, are provided. The Code Hub is a collection of open-source MRE software that includes a landing page with search functionality, linked to files hosted on the MRE Code Hub GitHub organization. The Code Catalog is a searchable online platform for discovery of useful (open-source or commercial) software packages, tools, codes, and other software products. To gather information about the existing MRE software landscape, a software survey is being performed, the preliminary results of which are presented herein. Initially, the data collected in the MRE software survey will be used to populate the MRE Software knowledge hub on PRIMRE, and future work will use data from the survey to perform a gap analysis and develop a vision for future software development. Additionally, as one of PRIMRE's roles is to support development of MRE software within project partners, a silo of knowledge relating to best practices has been gathered. An early draft of new guidance developed from this knowledge is presented.

gap analysis↗

Methodology for Thermodynamic Analysis Coupled with Computational Fluid Dynamics Modeling for Casting a Novel Aluminum–Cerium Alloy

In this paper, we present the results of computational fluid dynamics (CFD) analysis to assess castability, porosity, molten metal fluidity and other technological properties of a new Al-Ce alloy. A thermodynamic analysis of a new Al-Ce alloy is done to obtain the physical properties as a function of temperature across both solid and liquid phases. These properties are then used to build the CFD of model the full casting process from initial pouring through to final solidification of the part, here an example of a heavy duty mor mount is used. While several excellent commercial CFD codes exists this process shows that CFD can be used to assess the capability of the alloy to properly fill the mold, as well as give predictions where scattered porosity or large-scale defects may occur in the casting. Further, unlike the commercially available software (e.g., ProCAST, SOLIDCast , MAGMASOFT®) the complex 3D-analysis of the stress / strain fields in cast parts is not performed at this time. However, the availability of free software for assessing the required thermodynamic / thermo-physical properties of new alloys (OpenCALPHAD) and CFD codes such as OpenFOAM® makes the developed option attractive and economical, especially for the analysis of new Al-Ce alloys, for which the available data does not exist.

36 MATERIALS SCIENCE↗

A scalable matrix-free spectral element approach for unsteady PDE constrained optimization using PETSc/TAO

In this work, we provide a new approach for the efficient matrix-free application of the transpose of the Jacobian for the spectral element method for the adjoint-based solution of partial differential equation (PDE) constrained optimization. This results in optimizations of nonlinear PDEs using explicit integrators where the integration of the adjoint problem is not more expensive than the forward simulation. Solving PDE constrained optimization problems entails combining expertise from multiple areas, including simulation, computation of derivatives, and optimization. The Portable, Extensible Toolkit for Scientific computation (PETSc) together with its companion package, the Toolkit for Advanced Optimization (TAO), is an integrated numerical software library that contains an algorithmic/software stack for solving linear systems, nonlinear systems, ordinary differential equations, differential algebraic equations, and large-scale optimization problems and, as such, is an ideal tool for performing PDE-constrained optimization. This paper describes an efficient approach in which the software stack provided by PETSc/TAO can be used for large-scale nonlinear time-dependent problems. Time integration can involve a range of high-order methods, both implicit and explicit. The PDE-constrained optimization algorithm used is gradient-based and seamlessly integrated with the simulation of the physical problem.

97 MATHEMATICS AND COMPUTING↗

Distributed Macroscopic Traffic Simulation with Open Traffic Models

This paper presents OTM-MPI, an extension of the Open Traffic Models platform (OTM) for running macroscopic traffic simulations in high-performance computing environments. OTM-MPI represents the first open-source, distributed-memory, macroscopic simulation model developed for modern high performance parallel machines and large networks. Macroscopic simulations are appropriate for studying regional traffic scenarios when aggregate trends are of interest, rather than individual vehicle traces. They are also appropriate for studying the routing behavior of classes of vehicles, such as app-informed vehicles. The network partitioning was performed with METIS. Inter-process communication was done with MPI (message-passing interface). Results are provided for two networks: one realistic network which was obtained from Open Street Maps for Chattanooga, TN, and another larger synthetic grid network. The software recorded a speedups of 198x using 256 cores for Chattanooga, and 475x with 1,024 cores for the synthetic network.

macro-scopic traffic simulation↗