Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Taming the virtual space for incremental full configuration interaction

Incremental full configuration interaction (iFCI) closely approximates the FCI limit with polynomial cost through a many-body expansion of the correlation energy, providing highly accurate total energies within a given basis set. To extend iFCI beyond previous basis set limitations, this work introduces a novel natural orbital (NO) screening approach, incremental NO full configuration interaction (iNO-FCI). By consideration of the importance of virtual orbital selection in the convergence of iFCI, iNO-FCI maximizes the consistency between orbitals selected for each correlated body. iNO-FCI employs a principle of cancellation of errors and ensures that the same set of virtual NOs is used for interdependent terms. Here, this strategy significantly reduces computational cost without compromising precision. Computational savings of up to 95% are demonstrated, allowing access to larger basis sets that were previously computationally prohibitive. iNO-FCI is herein introduced and benchmarked for several difficult test cases involving double-bond dissociation, biradical systems, conjugated π systems, and the spin gap of a Cu-based transition metal complex.

Correlation energy↗

Tackling the Challenges in Scene Graph Generation With Local-to-Global Interactions

In this work, we seek new insights into the underlying challenges of the scene graph generation (SGG) task. Quantitative and qualitative analysis of the visual genome (VG) dataset implies: 1) ambiguity: even if interobject relationship contains the same object (or predicate), they may not be visually or semantically similar; 2) asymmetry: despite the nature of the relationship that embodied the direction, it was not well addressed in previous studies; and 3) higher-order contexts: leveraging the identities of certain graph elements can help generate accurate scene graphs. Motivated by the analysis, we design a novel SGG framework, Local-to-global interaction networks (LOGINs). Locally, interactions extract the essence between three instances of subject, object, and background, while baking direction awareness into the network by explicitly constraining the input order of subject and object. Globally, interactions encode the contexts between every graph component (i.e., nodes and edges). Finally, Attract and Repel loss is utilized to fine-tune the distribution of predicate embeddings. By design, our framework enables predicting the scene graph in a bottom-up manner, leveraging the possible complementariness. To quantify how much LOGIN is aware of relational direction, a new diagnostic task called Bidirectional Relationship Classification (BRC) is also proposed. Overall, experimental results demonstrate that LOGIN can successfully distinguish relational direction than existing methods (in BRC task), while showing state-of-the-art results on the VG benchmark (in SGG task).

97 MATHEMATICS AND COMPUTING↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Automated Generation of Message-Passing Programs: An Evaluation of CAPTools using NAS Benchmarks

Scientists at NASA Ames Research Center have been developing computational aeroscience applications on highly parallel architectures over the past ten years. During the same time period, a steady transition of hardware and system software also occurred, forcing us to expand great efforts into migrating and receding our applications. As applications and machine architectures continue to become increasingly complex, the cost and time required for this process will become prohibitive. Various attempts to exploit software tools to assist and automate the parallelization process have not produced favorable results. In this paper, we evaluate an interactive parallelization tool, CAPTools, for parallelizing serial versions of the NAB Parallel Benchmarks. Finally, we compare the performance of the resulting CAPTools generated code to the hand-coded benchmarks on the Origin 2000 and IBM SP2. Based on these results, a discussion on the feasibility of automated parallelization of aerospace applications is presented along with suggestions for future work.

Hribar, Michelle R.↗

Electra: A Modular-Based Expansion of NASA's Supercomputing Capability

NASA has increasingly relied on high-performance computing (HPC) re- sources for computational modeling, simulation, and data analysis to meet the science and engineering goals of its missions in space exploration, aeronautics, and Earth and space science. The NASA Advanced Supercomputing (NAS) Division at Ames Research Center in Silicon Valley, Calif., hosts NASA’s premier supercomputing resources, integral to achieving and enhancing the success of the agency’s missions. NAS provides a balanced environment, funded under the High-End Computing Capability (HECC) project, comprised of world-class supercomputers, including its flagship distributed-memory cluster, Pleiades; high-speed networking; and massive data storage facilities, along with multi-disciplinary support teams for user support, code porting and optimization, and large-scale data analysis and scientific visualization. However, as scientists have increased the fidelity of their simulations and engineers are conducting larger parameter-space studies, the requirements for supercomputing resources have been growing by leaps and bounds. With the facility housing the HECC systems reaching its power and cooling capacity, NAS undertook a prototype project to investigate an alternative approach for housing supercomputers. Modular supercomputing, or container-based computing, is an innovative concept for expanding NASA’s HPC capabilities. With modular supercomputing, additional containers—similar to portable storage pods—can be connected together as needed to accommodate the agency’s ever-increasing demand for computing resources. In addition, taking advantage of the local weather permits the use of cooling technologies that would additionally save energy and reduce annual water usage. The first stage of NASA’s Modular Supercomputing Facility (MSF) prototype, which resulted in a 1,000 square-foot module on a concrete pad with room for 16 compute racks, was completed in Fall 2016 and an SGI (now HPE) computer system, named Electra, was deployed there in early 2017. Cooling is performed via an evaporative system built into the module, and preliminary experience shows a Power Usage Effectiveness (PUE) measurement of 1.03. Electra achieved over a petaflop on the LINPACK benchmark, sufficient to rank number 96 on the November 2016 TOP500 list [14]. The system consists of 1,152 InfiniBand-connected Intel Xeon Broadwell-based nodes. Its users access their files on a facility-wide file system shared by all HECC compute assets via Mellanox MetroX InfiniBand extenders, which connect the Electra fabric to Lustre routers in the primary facility over fiber-optic links about 900 feet long. The MSF prototype has exceeded expectations and is serving as a blueprint for future expansions. In the remainder of this chapter, we detail how modular data center technology can be used to expand an existing compute resource. We begin by describing NASA’s requirements for supercomputing and how resources were provided prior to the integration of the Electra module-based system.

Biswas, Rupak↗

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking↗

Dynamic signatures of electronically nonadiabatic coupling in sodium hydride: a rigorous test for the symmetric quasi-classical model applied to realistic, ab initio electronic states in the adiabatic representation

Sodium hydride (NaH) in the gas phase presents a seemingly simple electronic structure making it a potentially tractable system for the detailed investigation of nonadiabatic molecular dynamics from both computational and experimental standpoints. The single vibrational degree of freedom, as well as the strong nonadiabatic coupling that arises from the excited electronic states taking on considerable ionic character, provides a realistic chemical system to test the accuracy of quasi-classical methods to model population dynamics where the results are directly comparable against quantum mechanical benchmarks. Here, using a simulated pump–probe type experiment, this work presents computational predictions of population transfer through the avoided crossings of NaH via symmetric quasi-classical Meyer–Miller (SQC/MM), Ehrenfest, and exact quantum dynamics on realistic, ab initio potential energy surfaces. The main driving force for population transfer arises from the ground vibrational level of the D 1 Σ + adiabatic state that is embedded in the manifold of near-dissociation C 1 Σ + vibrational states. When coupled through a sharply localized first-order derivative coupling most of the population transfers between t = 15 and t = 30 fs depending on the initially excited vibronic wavepacket. While quantum mechanical effects are expected due to the reduced mass of NaH, predictions of the population dynamics from both the SQC/MM and Ehrenfest models perform remarkably well against the quantum dynamics benchmark. Additionally, an analysis of the vibronic structure in the nonadiabatically coupled regime is presented using a variational eigensolver methodology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Anderson acceleration with approximate calculations: Applications to scientific computing

Here we provide rigorous theoretical bounds for Anderson acceleration (AA) that allow for approximate calculations when applied to solve linear problems. We show that, when the approximate calculations satisfy the provided error bounds, the convergence of AA is maintained while the computational time could be reduced. We also provide computable heuristic quantities, guided by the theoretical error bounds, which can be used to automate the tuning of accuracy while performing approximate calculations. For linear problems, the use of heuristics to monitor the error introduced by approximate calculations, combined with the check on monotonicity of the residual, ensures the convergence of the numerical scheme within a prescribed residual tolerance. Motivated by the theoretical studies, we propose a reduced variant of AA, which consists in projecting the least-squares used to compute the Anderson mixing onto a subspace of reduced dimension. The dimensionality of this subspace adapts dynamically at each iteration as prescribed by the computable heuristic quantities. We numerically show and assess the performance of AA with approximate calculations on: (i) linear deterministic fixed-point iterations arising from the Richardson's scheme to solve linear systems with open-source benchmark matrices with various preconditioners and (ii) non-linear deterministic fixed-point iterations arising from non-linear time-dependent Boltzmann equations.

97 MATHEMATICS AND COMPUTING↗

The multichannel i -propyl + O2 reaction system: A model of secondary alkyl radical oxidation

The i-propyl + O2 reaction mechanism has been investigated by definitive quantum chemical methods to establish this system as a benchmark for the combustion of secondary alkyl radicals. Focal point analyses extrapolating to the ab initio limit were performed based on explicit computations with electron correlation treatments through coupled cluster single, double, triple, and quadruple excitations and basis sets up to cc-pV5Z. The rigorous coupled cluster single, double, and triple excitations/cc-pVTZ level of theory was used to fully optimize all reaction species and transition states, thus, removing some substantial flaws in reference geometries existing in the literature. The vital i-propylperoxy radical (MIN1) and its concerted elimination transition state (TS1) were found 34.8 and 4.4 kcal mol−1 below the reactants, respectively. Two β-hydrogen transfer transition states (TS2, TS2′) lie above the reactants by (1.4, 2.5) kcal mol−1 and display large Born–Oppenheimer diagonal corrections indicative of nearby surface crossings. An α-hydrogen transfer transition state (TS5) is discovered 5.7 kcal mol−1 above the reactants that bifurcates into equivalent α-peroxy radical hanging wells (MIN3) prior to a highly exothermic dissociation into acetone + OH. The reverse TS5 → MIN1 intrinsic reaction path also displays fascinating features, including another bifurcation and a conical intersection of potential energy surfaces. An exhaustive conformational search of two hydroperoxypropyl (QOOH) intermediates (MIN2 and MIN3) of the i-propyl + O2 system located nine rotamers within 0.9 kcal mol−1 of the corresponding lowest-energy minima.

Chemistry↗

hypredrive: high-level interface for solving linear systems with hypre

This software introduces a high-level interface designed to simplify solving linear systems using hypre, a renowned library for such computational challenges. It is crafted to be accessible and user-friendly, making the powerful capabilities of hypre available to a broader audience without requiring in-depth technical knowledge. The interface is characterized by its use of YAML for input, a format celebrated for its structured yet straightforward readability. This choice ensures that users can easily configure the software to meet their specific needs. Additionally, the software boasts an intuitive API that encapsulates hypre's functionalities, making it easier for users to interact with the process of solving linear systems. It is particularly beneficial for prototyping, offering a quick and efficient means to test various solver and preconditioner configurations. Furthermore, the software allows for the creation of an offline testing framework in which predefined linear systems are read from files and benchmarked with user-defined solution strategies. This makes it an invaluable tool for developers and researchers exploring and validating their computational models. Overall, the software serves as a bridge, bringing the advanced computational capabilities of hypre closer to users who may need more specialized technical expertise, thereby facilitating innovation and exploration in the field of numerical linear algebra.

Paludetto Magri, Victor↗

Performance of BLAS 3, FFTs and NAS Parallel Benchmarks on Cray T3D

Recently, a Cray T3D Emulator has been made available on the Cray Y-MP and C90 computers. The Pittsburgh Supercomputer Center has acquired a CRAY T3D system and many other centers like Jet Propulsion Laboratory (JPL) will have it by the end of 1994. The Cray T3D system is the firstphase system in Cray Research, Inc.'s (CRI) three-phase massively parallel processing (MPP) program. This system features a heterogeneous architecture that closely couples DEC's ALPHA microprocessors and CRI's parallel-vector technology, i.e. the Cray Y-MP and Cray C90. The Cray T3D Emulator will give prospective users a valuable experience in developing high performance applications on the MPP system. This emulator runs programs written in CRI's MPP Fortran programming model (data sharing and work sharing) or Parallel Virtual Machine (PVM) programming model. It will help the users to study data layout, data locality, and data reference patterns thereby providing feedback which will enable one to write more efficient parallel codes. An overview of the Cray T3D hardware, software, and three of its available programming models is presented.The Cray Fortran Programming Model comprising (a) Data Sharing, (b) Worksharing and (c) Message Passing, will be discussed with examples. We have also implemented distributed BLAS 3 (matrix-matrix multiplication) in data parallel model (using only CSHIFT); worksharing model using block distribution and collapsed distribution; and message passing model using PVM. We have also implemented 2D and 3D FFTs for radix-2 using PVM. The performance of NAS Parallel 'Benchmarks (NPB) on CRAY T3D will be compared with other highly parallel systems such as CM-5, Paragon, C90 etc.

Saini, Subhash↗

Correlating AGP on a quantum computer

For variational algorithms on the near term quantum computing hardware, it is highly desirable to use very accurate ansatze with low implementation cost. Recent studies have shown that the antisymmetrized geminal power (AGP) wavefunction can be an excellent starting point for ansatze describing systems with strong pairing correlations, as those occurring in superconductors. In this work, we show how AGP can be efficiently implemented on a quantum computer with circuit depth, number of CNOTs, and number of measurements being linear in system size. Using AGP as the initial reference, we propose and implement a unitary correlator on AGP and benchmark it on the ground state of the pairing Hamiltonian. Furthermore, the results show highly accurate ground state energies in all correlation regimes of this model Hamiltonian.

97 MATHEMATICS AND COMPUTING↗

Certifying almost all quantum states with few single-qubit measurements

Certifying that an n -qubit state synthesized in the laboratory is close to a given target state is a fundamental task in quantum information science. However, existing rigorous protocols applicable to general target states have potentially prohibitive resource requirements in the form of either deep quantum circuits or exponentially many single-qubit measurements. Here we prove that almost all n -qubit target states, including those with exponential circuit complexity, can be certified from only O ( n 2 ) single-qubit measurements. Given access to the target state’s amplitudes, our protocol requires only O ( n 3 ) classical computation. This result is established by a technique that relates certification to the mixing time of a random walk. Our protocol has applications for benchmarking quantum systems, for optimizing quantum circuits to generate a desired target state and for learning and verifying neural networks, tensor networks and various other representations of quantum states using only single-qubit measurements. We show that such verified representations can be used to efficiently predict highly non-local properties of a synthesized state that would otherwise require an exponential number of measurements on the state. We demonstrate these applications in numerical experiments with up to 120 qubits and observe an advantage over existing methods such as cross-entropy benchmarking.

information theory and computation↗

System Noise Benchmarks

This project includes benchmarks to assess the presence of system noise on supercomputers. System noise is any activity that interferes with the execution of high-performance computing applications.

Moody, AdamT [Lawrence Livermore National Laborato↗

Verified, Archived, Library of Inputs and Data (VALID) Supporting Files

This dataset contains input, output, and sensitivity data files for computational simulations with the SCALE code system as part of the Verified, Archived Library of Inputs and Data (VALID). The simulations cover critical benchmark experiments from the International Criticality Safety Benchmark Evaluation Project. The files are to be housed in a public directory for distribution. The information contained in the files have been approved for release by the Organisation for Economic Co-operation and Development Nuclear Energy Agency (NEA). Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases.

keff↗

Particle Swarm Optimization Algorithm for Critical Experiment Design

Nuclear criticality experiments are used to validate nuclear cross section data used by simulation software. This is typically achieved by designing a critical system with a high sensitivity to a certain material’s cross section. Once the experiment has been carried out, a high fidelity model of the system is developed into a benchmark. When this benchmark model is simulated by a transport code, some of the difference between the experimental and computational effective neutron multiplication factor can be attributed to inaccurate nuclear data. Nuclear data evaluators then can make adjustments accordingly to improve cross section data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Fan Noise Prediction with Applications to Aircraft System Noise Assessment

This paper describes an assessment of current fan noise prediction tools by comparing measured and predicted sideline acoustic levels from a benchmark fan noise wind tunnel test. Specifically, an empirical method and newly developed coupled computational approach are utilized to predict aft fan noise for a benchmark test configuration. Comparisons with sideline noise measurements are performed to assess the relative merits of the two approaches. The study identifies issues entailed in coupling the source and propagation codes, as well as provides insight into the capabilities of the tools in predicting the fan noise source and subsequent propagation and radiation. In contrast to the empirical method, the new coupled computational approach provides the ability to investigate acoustic near-field effects. The potential benefits/costs of these new methods are also compared with the existing capabilities in a current aircraft noise system prediction tool. The knowledge gained in this work provides a basis for improved fan source specification in overall aircraft system noise studies.

Nark, Douglas M.↗

Modeling of the Molten Salt Reactor Experiment with SCALE

A SCALE model was developed for the Molten Salt Reactor Experiment (MSRE) benchmark that was recently added to the International Handbook of Evaluated Reactor Physics Benchmark Experiments. This SCALE model served as a basis for criticality calculations and nuclear data sensitivity and uncertainty analyses with the Monte Carlo code Shift and the TSUNAMI computational capabilities in the SCALE code system. The focus of this work is the assessment of the impact of nuclear data on the calculated eigenvalue results in support of the discussion of differences between the calculated and the experimental eigenvalue result. The differences in the eigenvalues obtained using the ENDF/B-VII.0, ENDF/B-VII.1, and ENDF/B-VIII.0 nuclear data libraries cover a relatively small range of ~230 pcm. Since eigenvalue sensitivity of the MSRE is dominated by the neutron multiplicity and neutron capture of 235 U and elastic scattering in graphite, relevant changes in the ENDF/B libraries for nuclear reactions (such as carbon capture) that caused large differences in other graphite-moderated systems did not have a significant impact. Propagation of nuclear data uncertainty results in an eigenvalue uncertainty of ~700 pcm with the major contributors being 235 U neutron multiplicity, graphite elastic scattering, and 7Li neutron capture. All calculations resulted in large differences of ~2000 pcm in eigenvalue compared to the benchmark experimental value. Several potential contributors to this difference—including uncertainties and gaps in the knowledge of the material, geometry, and nuclear data—were identified. Simplified models of the full MSRE core were developed, and similarity assessments were conduced with the full MSRE core model. It was found that simplified models can serve as adequate surrogates of the full-core model such that they can be used for performing selected nuclear data performance assessments with a lower computational burden.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗