Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “performance portable algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Development of NDE/NDT Tools for High-Volume & High-Speed Inspection of CFRP Structures in Automotive Manufacturing

Main advantages of the air-coupled ultrasound testing (ACUT) and electromagnetic testing (EMT) techniques for NDE of CFRP composites were non-contact sensing, scalability for high-speed inspection, cost-effectiveness, and non-hazardous operation. Despite these advantages, no systems that would satisfy the project requirements were commercially available. Hence, one of the major efforts of the Michigan State University (MSU) team at the initial stage of the project was to close this technological gap by developing, optimizing, and validating array sensors that would provide sufficient sensitivity, spatial coverage, and resolution for robust defect detection. Optimization of the ACUT and EMT sensor designs was performed using experimentally validated finite element models. Initial experiments using array probes were conducted on relatively flat CFRP samples. In parallel, the MSU team designed and assembled a portable platform with two robotic arms. The robots were equipped with newly designed sensors that enabled high-speed NDE of curved CFRP parts. Presently, the developed robotic platform can be used as a demo/template NDE system, which is easily adaptable to manufacturing environments and in-line NDE. The ACUT NDE system developed by the MSU team used a high-power 4-channel pulser receiver for parallel data acquisition. The array probes were designed by stacking commercially available ACUT transducers, which operated in the frequency range between 100 kHz and 500 kHz. MSU optimized the excitation procedure and developed wave focusing cones so as to reduce the crosstalk between the transducers and to provide higher pulse repletion frequency (PRF). The through-transmission (TT) and single-side access (SSA) inspection modes were successfully implemented. In the TT-ACUT, structural defects in CFRP were detected by passing ultrasonic waves through the test part. Hence, the ACUT transmitters and receivers needed to be placed on the opposite sides of the test part. In the SSA-ACUT, guided waves (GW) were excited in the test part using the transmitters and were sensed by the receivers from the same side. Multi-channel TT-ACUT and SSA-ACUT provided high-speed NDE, and were successfully validated on CFRP test samples with interlaminar delaminations and other embedded defects The EM techniques developed by the MSU team included: 1) eddy current testing (ECT), 2) capacitive imaging (CI) and hybrid dual-mode imaging. In ECT, structural damage was detected in CFRP using coils sensor arrays. In ECT, the excitation magnetic field is generated by passing an alternating current through a coil, which is placed above the test sample. The excitation field penetrates the conductive sample and induces the eddy currents in its transect. In turn, the eddy currents generate the reaction field, which affects the total field sensed by a coil. Hence, the presence of structural flaws will alter the eddy current flow and the picked-up signal. ECT is mostly sensitive to local changes of the electric conductivity of the test sample, and CFRPs are mostly conductive in the direction of carbon fibers. Hence, ECT was well suited for the detection of fiber damage/fiber irregularities. The MSU team developed printed circuit boards (PCB) with coil sensor arrays optimized for NDE of CFRP. Unlike most commercial probes designed for ECT of metallic structures, the MSU array probes were designed for operation in [1-10] MHz frequency range, which was optimal for low-conductive CFRP. Multiple sensing topologies (coil groups excitation/sensing arrangements) were implemented and successfully validated. Capacitive Imaging (CI) technique developed by MSU was complementary to ECT. In contrast to ECT, which was sensitive to local changes of the electrical conductivity, the CI was sensitive to local changes of the dielectric constant. Therefore, CI could provide information about matrix damage/matrix irregularities in CFRP. The MSU CI sensor arrays were made of multiple circular or rectangular open-plate capacitors printed on PCB. Sensors of this type are not commercially available. In addition to ECT and CI, the MSU team developed a hybrid (dual-mode) inductive/capacitive measurement technique that synergistically combined the benefits of inductive and capacitive sensing for rapid NDE of fiber reinforced polymer (FRP) composite structures. Fiber damage and fiber irregularities in FRPs were detected by configuring hybrid sensors as coil sensors. Similarly, matrix damage, matrix irregularities and interlaminar delaminations were detected by configuring hybrid sensors as capacitive sensors. ECT and CI were performed sequentially by means of electronic switching. Hence, eliminating the need for mounting two separate sensor arrays on the probe. Portable robotic platform was developed by MSU for multi-technique high-speed NDE of CFRP test parts. The platform had two 6-axis robots, which enabled inspection of curved parts in approximately a 6×6×6 ft 3 active scan area. On the software side, the MSU team integrated scripts for NDE hardware control with scripts for robot motion control. MSU also implemented automated path planning for the robots, reconstruction of part’s surfaces via stereovision, 3D rendering of inspection data, and image processing algorithms for enhanced defect detection. Automotive composite parts manufactured by Plasan Composites from Phase I were used to validate the ACUT and EMT techniques on representative testbeds. Among those parts were three X-braces for a Dodge Viper, one composite calibration plaque with known defects at known locations, and four other test sections, including sections from a front splitter, a corner section from a composite hood, and a high-pressure RTM panel made using non crimp fabric. Other test samples included CFRP and GFRP calibration plates with fiber/matrix defects fabricated at MSU/CVRC.

36 MATERIALS SCIENCE↗

A New Electromagnetic Instrument for Thickness Gauging of Conductive Materials

Eddy current techniques are widely used to measure the thickness of electrically conducting materials. The approach, however, requires an extensive set of calibration standards and can be quite time consuming to set up and perform. Recently, an electromagnetic sensor was developed which eliminates the need for impedance measurements. The ability to monitor the magnitude of a voltage output independent of the phase enables the use of extremely simple instrumentation. Using this new sensor a portable hand-held instrument was developed. The device makes single point measurements of the thickness of nonferromagnetic conductive materials. The technique utilized by this instrument requires calibration with two samples of known thicknesses that are representative of the upper and lower thickness values to be measured. The accuracy of the instrument depends upon the calibration range, with a larger range giving a larger error. The measured thicknesses are typically within 2-3% of the calibration range (the difference between the thin and thick sample) of their actual values. In this paper the design, operational and performance characteristics of the instrument along with a detailed description of the thickness gauging algorithm used in the device are presented.

Fulton, J. P.↗

Extending SEER for Extreme Heterogeneity

Heterogeneous and multi-device nodes are increasingly common in high-performance computing and data centers, yet existing programming models often lack simple, transparent, and portable support for these diverse architectures. The main contribution of this work is the development of novel SEER capabilities to address this challenge by providing a descriptive programming model that allows applications to seamlessly leverage heterogeneous nodes across various device types. SEER uses efficient memory management and can select the proper device[s] depending on the computational cost of the applications. This is completely transparent to the programmer, thereby providing a highly productive programming environment. Integrating extreme heterogeneity into the SEER library as shown with the use of NVIDIA and AMD GPUs simultaneously allows it to expand and exploit the performance possibilities. Our analysis based on the well-known Conjugate Gradient algorithm reports accelerations above 1.5 × on computationally demanding steps of such an algorithm by using both architectures simultaneously.

Teranishi, Keita [ORNL] (ORCID:0000000166472690)↗

Load Balancing Sequences of Unstructured Adaptive Grids

Mesh adaption is a powerful tool for efficient unstructured grid computations but causes load imbalance on multiprocessor systems. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. This paper makes several important additions to our previous work. First, a new remapping cost model is presented and empirically validated on an SP2. Next, our load balancing strategy is applied to sequences of dynamically adapted unstructured grids. Results indicate that our framework is effective on many processors for both steady and unsteady problems with several levels of adaption. Additionally, we demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required for a fine initial mesh. Finally, we show that the data remapping overhead can be significantly reduced by applying our heuristic processor reassignment algorithm.

Biswas, Rupak↗

Application of Gaussian Bayes classifier to differentiate chlorine-based chemical agents

The Portable Isotopic Neutron Spectroscopy (PINS) is a commercialized system developed by Idaho National Laboratory (INL) to examine chemical warfare agents (CWA) non-destructively, utilizing Prompt Gamma Neutron Activation Analysis (PGNAA) techniques. The PINS system takes advantage of a high-resolution gamma-ray spectrum from a mechanically-cooled high-purity germanium (HPGe) detector. One of the difficult technical challenges is to discriminate the chlorine-based chemical agents. Especially, CN, CNB, CNS and CG have similar chemical compositions to make it hard to discriminate them with a higher confidence. Current identification algorithms for PINS systems with 252-Cf sources have been improved and updated continuously as more field data became available, and new algorithms was studied to complement the current algorithms by adopting the Gaussian Bayes classifier. These new algorithms were intended to be applied to a subset of chlorine-based chemical agents, and their main goal is discriminate CN, CNB, CNS and CG with their ratios of the chlorine neutron inelastic 1763keV peak to the chlorine thermal neutron capture 1959keV peak, which is referred to as the “Cl i/c” or “clic” ratio in this study. The Cl i/c ratios were assumed to follow Gaussian distributions with the means and the standard deviations unique to their corresponding chemical agents. The prior probabilities of these four chemical agents were optimized with a collection of field data to achieve the best performance in terms of precision or positive predictive value (PPV). Finally, their posterior probabilities as functions of the Cl i/c ratio were implemented in the current version of PINS analysis software in order to be tested with more field data.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

NASA Tech Briefs, November 2007

Topics include: Wireless Measurement of Contact and Motion Between Contact Surfaces; Wireless Measurement of Rotation and Displacement Rate; Portable Microleak-Detection System; Free-to-Roll Testing of Airplane Models in Wind Tunnels; Cryogenic Shrouds for Testing Thermal-Insulation Panels; Optoelectronic System Measures Distances to Multiple Targets; Tachometers Derived From a Brushless DC Motor; Algorithm-Based Fault Tolerance for Numerical Subroutines; Computational Support for Technology- Investment Decisions; DSN Resource Scheduling; Distributed Operations Planning; Phase-Oriented Gear Systems; Freeze Tape Casting of Functionally Graded Porous Ceramics; Electrophoretic Deposition on Porous Non- Conductors; Two Devices for Removing Sludge From Bioreactor Wastewater; Portable Unit for Metabolic Analysis; Flash Diffusivity Technique Applied to Individual Fibers; System for Thermal Imaging of Hot Moving Objects; Large Solar-Rejection Filter; Improved Readout Scheme for SQUID-Based Thermometry; Error Rates and Channel Capacities in Multipulse PPM; Two Mathematical Models of Nonlinear Vibrations; Simpler Adaptive Selection of Golomb Power-of- Two Codes; VCO PLL Frequency Synthesizers for Spacecraft Transponders; Wide Tuning Capability for Spacecraft Transponders; Adaptive Deadband Synchronization for a Spacecraft Formation; Analysis of Performance of Stereoscopic-Vision Software; Estimating the Inertia Matrix of a Spacecraft; Spatial Coverage Planning for Exploration Robots; and Increasing the Life of a Xenon-Ion Spacecraft Thruster.

Source record↗

Lifting the Garage Door on Spawn, An Open-Source BEM-Controls Engine

Spawn is the latest whole-building energy simulation engine developed by the US Department of Energy, National Labs and industry. Whereas EnergyPlus was designed as a successor to DOE-2, Spawn is not a direct successor of–nor is it intended as an imminent replacement for– EnergyPlus. Instead, Spawn reuses parts of EnergyPlus while supporting new use cases in HVAC and controls. Spawn is intended to provide several capabilities that significantly advance beyond EnergyPlus. It is intended to support the evaluation of novel HVAC and district energy systems in a more physically realistic way. Critically, it can model control in a physically realistic way, using portable specifications that can be compiled for execution on control platforms. Spawn is also intended to support co-simulation in an intrinsic way to enable integration with third-party models. This paper describes the software architecture of Spawn from model authoring to compilation and simulation. It explains how Spawn reuses the envelope and daylighting modules of EnergyPlus and couples them to HVAC and control models from the Modelica Buildings Library using the Functional Mockup Interface (FMI) standard. It presents a number of examples that: i) validate Spawn’s coupled simulation approach by comparing its results to those of EnergyPlus, ii) illustrate the Spawn methodology for modeling and simulating HVAC systems, and iii) evaluate the performance of Spawn’s Quantized State System (QSS) time integration algorithms

Wetter, Michael↗

Enabling Multireference Calculations on Multimetallic Systems with Graphic Processing Units

Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and the design of new materials. This work describes an initial implementation for performing multireference wave function method localized active-space self-consistent field (LASSCF) calculations through the use of multiple graphics processing units (GPUs) to accelerate time-to-solution. Density fitting is leveraged to reduce memory requirements, and we demonstrate the ability to fully utilize multi-GPU compute nodes. Performance improvements of 5–10x in total application runtime were observed in LASSCF calculations for multimetallic catalyst systems up to 1200 AOs and an active space of (22e,40o) using up to four NVIDIA A100 GPUs. Furthermore, written with performance portability in mind, a comparable performance is also observed in early runs on the Aurora exascale system using Intel Max Series GPUs.

Algorithms↗

Efficient phase-space generation for hadron collider event simulation

We present a simple yet efficient algorithm for phase-space integration at hadron colliders. Individual mappings consist of a single t-channel combined with any number of s-channel decays, and are constructed using diagrammatic information. The factorial growth in the number of channels is tamed by providing an option to limit the number of s-channel topologies. We provide a publicly available, parallelized code in C++ and test its performance in typical LHC scenarios.

47 OTHER INSTRUMENTATION↗

Sampling Technique for Robust Odorant Detection Based on MIT RealNose Data

This technique enhances the detection capability of the autonomous Real-Nose system from MIT to detect odorants and their concentrations in noisy and transient environments. The lowcost, portable system with low power consumption will operate at high speed and is suited for unmanned and remotely operated long-life applications. A deterministic mathematical model was developed to detect odorants and calculate their concentration in noisy environments. Real data from MIT's NanoNose was examined, from which a signal conditioning technique was proposed to enable robust odorant detection for the RealNose system. Its sensitivity can reach to sub-part-per-billion (sub-ppb). A Space Invariant Independent Component Analysis (SPICA) algorithm was developed to deal with non-linear mixing that is an over-complete case, and it is used as a preprocessing step to recover the original odorant sources for detection. This approach, combined with the Cascade Error Projection (CEP) Neural Network algorithm, was used to perform odorant identification. Signal conditioning is used to identify potential processing windows to enable robust detection for autonomous systems. So far, the software has been developed and evaluated with current data sets provided by the MIT team. However, continuous data streams are made available where even the occurrence of a new odorant is unannounced and needs to be noticed by the system autonomously before its unambiguous detection. The challenge for the software is to be able to separate the potential valid signal from the odorant and from the noisy transition region when the odorant is just introduced.

Duong, Tuan A.↗

Tensor Network Quantum Virtual Machine for Simulating Quantum Circuits at Exascale

The numerical simulation of quantum circuits is an indispensable tool for development, verification, and validation of hybrid quantum-classical algorithms intended for near-term quantum co-processors. The emergence of exascale high-performance computing (HPC) platforms presents new opportunities for pushing the boundaries of quantum circuit simulation. Here, we present a modernized version of the Tensor Network Quantum Virtual Machine (TNQVM) that serves as the quantum circuit simulation backend in the eXtreme-scale ACCelerator (XACC) framework. The new version is based on the scalable tensor network processing library ExaTN (Exascale Tensor Networks). It provides multiple configurable quantum circuit simulators that perform either an exact quantum circuit simulation via the full tensor network contraction or an approximate simulation via a suitably chosen tensor factorization scheme. Upon necessity, stochastic noise modeling from real quantum processors is incorporated into the simulations by modeling quantum channels with Kraus tensors. By combining the portable XACC quantum programming frontend and the scalable ExaTN numerical processing backend, we introduce an end-to-end virtual quantum development environment that can scale from laptops to future exascale platforms. We report initial benchmarks of our framework, which include a demonstration of the distributed execution, incorporation of quantum decoherence models, and simulation of the random quantum circuits used for the certification of quantum supremacy on Google’s Sycamore superconducting architecture.

Nguyen, Thien↗

Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex

Functional near infrared spectroscopy (fNIRS) is an emerging optical neuroimaging technology that indirectly measures neuronal activity in the cortex via neurovascular coupling. It quantifies hemoglobin concentration ([Hb]) and thus measures the same hemodynamic response as functional magnetic resonance imaging (fMRI), but is portable, non-confining, relatively inexpensive, and is appropriate for long-duration monitoring and use at the bedside. Like fMRI, it is noninvasive and safe for repeated measurements. Patterns of [Hb] changes are used to classify cognitive state. Thus, fNIRS technology offers much potential for application in operational contexts. For instance, the use of fNIRS to detect the mental state of commercial aircraft operators in near real time could allow intelligent flight decks of the future to optimally support human performance in the interest of safety by responding to hazardous mental states of the operator. However, many opportunities remain for improving robustness and reliability. It is desirable to reduce the impact of motion and poor optical coupling of probes to the skin. Such artifacts degrade signal quality and thus cognitive state classification accuracy. Field application calls for further development of algorithms and filters for the automation of bad channel detection and dynamic artifact removal. This work introduces a novel adaptive filter method for automated real-time fNIRS signal quality detection and improvement. The output signal (after filtering) will have had contributions from motion and poor coupling reduced or removed, thus leaving a signal more indicative of changes due to hemodynamic brain activations of interest. Cognitive state classifications based on these signals reflect brain activity more reliably. The filter has been tested successfully with both synthetic and real human subject data, and requires no auxiliary measurement. This method could be implemented as a real-time filtering option or bad channel rejection feature of software used with frequency domain fNIRS instruments for signal acquisition and processing. Use of this method could improve the reliability of any operational or real-world application of fNIRS in which motion is an inherent part of the functional task of interest. Other optical diagnostic techniques (e.g., for NIR medical diagnosis) also may benefit from the reduction of probe motion artifact during any use in which motion avoidance would be impractical or limit usability.

Harrivel, Angela↗

Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex

Functional near infrared spectroscopy (fNIRS) is an emerging optical neuroimaging technology that indirectly measures neuronal activity in the cortex via neurovascular coupling. It quantifies hemoglobin concentration ([Hb]) and thus measures the same hemodynamic response as functional magnetic resonance imaging (fMRI), but is portable, non-confining, relatively inexpensive, and is appropriate for long-duration monitoring and use at the bedside. Like fMRI, it is noninvasive and safe for repeated measurements. Patterns of [Hb] changes are used to classify cognitive state. Thus, fNIRS technology offers much potential for application in operational contexts. For instance, the use of fNIRS to detect the mental state of commercial aircraft operators in near real time could allow intelligent flight decks of the future to optimally support human performance in the interest of safety by responding to hazardous mental states of the operator. However, many opportunities remain for improving robustness and reliability. It is desirable to reduce the impact of motion and poor optical coupling of probes to the skin. Such artifacts degrade signal quality and thus cognitive state classification accuracy. Field application calls for further development of algorithms and filters for the automation of bad channel detection and dynamic artifact removal. This work introduces a novel adaptive filter method for automated real-time fNIRS signal quality detection and improvement. The output signal (after filtering) will have had contributions from motion and poor coupling reduced or removed, thus leaving a signal more indicative of changes due to hemodynamic brain activations of interest. Cognitive state classifications based on these signals reflect brain activity more reliably. The filter has been tested successfully with both synthetic and real human subject data, and requires no auxiliary measurement. This method could be implemented as a real-time filtering option or bad channel rejection feature of software used with frequency domain fNIRS instruments for signal acquisition and processing. Use of this method could improve the reliability of any operational or real-world application of fNIRS in which motion is an inherent part of the functional task of interest. Other optical diagnostic techniques (e.g., for NIR medical diagnosis) also may benefit from the reduction of probe motion artifact during any use in which motion avoidance would be impractical or limit usability.

Harrivel, Angela↗

A comparative study on deep learning models for condition monitoring of advanced reactor piping systems

Advanced nuclear reactors offer innovative applications due to their portability, reliability, resiliency, and high capacity factors. To operate them on a wider scale, reducing maintenance life-cycle costs while ensuring their integrity is essential. Autonomous operations in advanced nuclear reactors using augmented Digital Twin (DT) technology can serve as a cost-effective solution by increasing awareness about the system’s health. A key component of nuclear DT frameworks is the condition monitoring of safety systems, such as piping-equipment systems, which involves acquiring and monitoring the plant’s sensor data. Here, this research proposes a condition monitoring methodology utilizing deep learning algorithms, such as multilayer perceptions (MLP) and convolutional neural networks (CNNs), to detect degradation and its severity in nuclear piping-equipment systems. Sensor signals are processed to obtain the power spectral density and the Short-Time Fourier transform, and feature extraction methodologies are proposed to develop degradation-sensitive data repositories. The performance of MLP, one-dimensional (1D) CNN, and 2D CNN within the proposed condition monitoring framework is compared using a finite element model of a 3D piping system subjected to seismic loads as the application case study. Various approaches, such as dropout, k-Fold validation, regularization, and early stopping of training the network, are investigated to avoid overfitting the models to the input sensor data. The predictive capability and computational capacity of the deep learning algorithms are also compared to detect degradation in the Z-pipe system of the Experimental Breeder Reactor II (EBRII). The Z-pipe system is subjected to harmonic excitations that represent normal operating loads, such as pump-induced vibrations. The findings of the study indicate that the proposed artificial intelligence (AI)-driven condition monitoring framework demonstrates superior prediction accuracies with a 2D CNN, whereas the MLP exhibits higher computational efficiency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 kin or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed- shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

Automated clustering-based workload characterization

The demands placed on the mass storage systems at various federal agencies and national laboratories are continuously increasing in intensity. This forces system managers to constantly monitor the system, evaluate the demand placed on it, and tune it appropriately using either heuristics based on experience or analytic models. Performance models require an accurate workload characterization. This can be a laborious and time consuming process. It became evident from our experience that a tool is necessary to automate the workload characterization process. This paper presents the design and discusses the implementation of a tool for workload characterization of mass storage systems. The main features of the tool discussed here are: (1)Automatic support for peak-period determination. Histograms of system activity are generated and presented to the user for peak-period determination; (2) Automatic clustering analysis. The data collected from the mass storage system logs is clustered using clustering algorithms and tightness measures to limit the number of generated clusters; (3) Reporting of varied file statistics. The tool computes several statistics on file sizes such as average, standard deviation, minimum, maximum, frequency, as well as average transfer time. These statistics are given on a per cluster basis; (4) Portability. The tool can easily be used to characterize the workload in mass storage systems of different vendors. The user needs to specify through a simple log description language how the a specific log should be interpreted. The rest of this paper is organized as follows. Section two presents basic concepts in workload characterization as they apply to mass storage systems. Section three describes clustering algorithms and tightness measures. The following section presents the architecture of the tool. Section five presents some results of workload characterization using the tool.Finally, section six presents some concluding remarks.

Pentakalos, Odysseas I.↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 km or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed-shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

Implementation of McMurchie–Davidson Algorithm for Gaussian AO Integrals Suited for SIMD Processors

We report an implementation of the McMurchie− Davidson evaluation scheme for 1- and 2-particle Gaussian AO integrals designed for processors with Single Instruction Multiple Data (SIMD) instruction sets. Like in our recent MD implementation for graphical processing units (GPUs) [Asadchev, A.; Valeev, E. F.. J. Chem. Phys. 2024, 160, 244109.], variable-sized batches of shellsets of integrals are evaluated at a time. By optimizing for the floating point instruction throughput rather than minimizing the number of operations, this approach achieves up to 50% of the theoretical hardware peak FP64 performance for many common SIMD-equipped platforms (AVX2, AVX512, NEON), which translates to speedups of up to 30 over the state-of-the-art one-shellset-at-a-time implementation of Obara−Saika-type schemes in Libint for a variety of primitive and contracted integrals. As with our previous work, we rely on the standard C++ programming language such as the std::simd standard library feature to be included in the 2026 ISO C++ standard without any explicit code generation to keep the code base small and portable. The implementation is part of the open source LibintX library freely available at https://github.com/ValeevGroup/libintx.

Basis sets↗