Nonlinear analysis and performance of energy harvesting absorbers with stoppers for controlling fluid-induced vibrations of dynamical systems.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Nuclear power continues to have an imperative role for the U.S.’s electricity generation. However, for these nuclear power plants (NPPs) to remain economically viable, new strategies for reducing operations and maintenance costs must be explored. The U.S. Department of Energy Light Water Reactor Sustainability Program is researching how to digitally transform existing analog and hybrid main control rooms into a fully integrated control room that addresses this challenge. The transformation will fundamentally change the conduct of operations for these U.S. NPPs. Human factors engineering has a vital role in this effort, where traditional task analysis methods are important in informing the design of advanced human-system interface displays. This work describes a preliminary tool to support task analysis for the development of these advanced displays. Details on the use of this tool, including the specific task analysis methods offered, are presented in this paper.
To address the challenge of performance analysis on the US DOE’s forthcoming exascale supercomputers, Rice University has been extending its HPCToolkit performance tools to support measurement and analysis of GPU-accelerated applications. To help developers understand the performance of accelerated applications as a whole, HPCToolkit’s measurement and analysis tools attribute metrics to calling contexts that span both CPUs and GPUs. To measure GPU-accelerated applications efficiently, HPCToolkit employs a novel wait-free data structure to coordinate monitoring and attribution of GPU performance. To help developers understand the performance of complex GPU code generated from high-level programming models, HPCToolkit constructs sophisticated approximations of call path profiles for GPU computations. To support fine-grained analysis and tuning, HPCToolkit uses PC sampling and instrumentation to measure and attribute GPU performance metrics to source lines, loops, and inlined code. To supplement fine-grained measurements, HPCToolkit can measure GPU kernel executions using hardware performance counters. To provide a view of how an execution evolves over time, HPCToolkit can collect, analyze, and visualize call path traces within and across nodes. Finally, on NVIDIA GPUs, HPCToolkit can derive and attribute a collection of useful performance metrics based on measurements using GPU PC samples. Here, we illustrate HPCToolkit’s new capabilities for analyzing GPU-accelerated applications with several codes developed as part of the Exascale Computing Project.
In this study, the Kokkos based library Cabana, which has been developed in the Co-design Center for Particle Applications (CoPA), is used for the implementation of Multi-Particle Collision Dynamics (MPCD), a particle-based description of hydrodynamic interactions. Cabana allows for a function portable implementation, which has been used to study the interplay between CPU and GPU usage on a multi-node system as well as analysis of said interplay with performance analysis tools. As a result, we see most advantages in a homogeneous GPU usage, but we also discuss the extent to which heterogeneous applications might be more performant, using both CPU and GPU concurrently.
The U.S. Department of Energy (DOE) operates a low-level radioactive waste (LLW) disposal site at Material Disposal Area G, in Los Alamos, New Mexico, USA. Area G has been the primary LLW disposal site for Los Alamos National Laboratory (LANL) since the 1960's. In addition to LLW, Area G is host to a variety of other wastes, the disposition of which must be determined before closure of the site. A probabilistic Radiological Risk Assessment (RRA) for Area G is used in order to support decision making regarding some wastes that are not addressed in the extant Area G Performance Assessment (PA) and Composite Analysis (CA). Between 1979 and 1987, 33 special shafts were augered into the Bandelier Tuff at Area G. This volcanic tuff is present across Pajarito Plateau on the eastern slopes of the Jemez Mountains, and varies widely in its consistency, from weakly indurated non-welded layers to welded layers that uphold the mesa cliffs of the Plateau. These mesas are home to LANL, Area G, and the townsites of Los Alamos and White Rock, with residences about 1400 m from Area G. The 33 Shafts were lined with steel casing, and contain remote-handled (RH) transuranic wastes (TRU) resulting from experiments and analysis performed in special glove boxes at the Chemistry and Metallurgy Research (CMR) facility at LANL. Some of these wastes originated as used nuclear fuel. The purpose of the Area G RRA is to evaluate the potential future risk to humans and the environment from the RH TRU in the 33 Shafts in the context of the risk associated with the surrounding wastes at Area G. The analysis is responsive to expectations outlined in DOE Order 458.1, Radiation Protection of the Public and the Environment, and is informed by the Manual and Guidance accompanying DOE O 435.1, Radioactive Waste Management. Because the waste meets the definition of TRU, the regulatory context necessarily takes into consideration the regulation governing the disposal of TRU from the U.S. Environmental Protection Agency (EPA): 40 CFR 191, Environmental Radiation Protection Standards for Management and Disposal of Spent Nuclear Fuel, High-Level and Transuranic Radioactive Wastes. Given the broader regulatory context for the RRA, the analysis is subject to different assumptions from those made in the existing DOE O 435.1 PA and CA, such as allowing for future occupation of the site. The analysis begins with a comprehensive evaluation of features, events, processes, and exposure scenarios (FEPS) for Area G and the wastes it contains. These FEPSs are screened to eliminate from further consideration those of extremely low probability and/or consequence, and a conceptual site model (CSM) is subsequently developed. The scope and structure of the Area G RRA Model is informed by this CSM, and the Area G RRA Model is developed using the GoldSim systems analysis modeling platform. This paper presents the initial version of a defensible, transparent, and reasonably realistic model, which is based on the state of knowledge of the wastes, the site, and the FEPSs that govern contaminant transport from wastes into the environment and subsequent exposures to humans and other biota. Probabilistic model input distributions represent uncertainties inherent in the real and modeled systems. The results of the Area G RRA Model inform decisions regarding the disposition of the RH TRU in the 33 Shafts. (authors)
High-resolution simulations of polar ice sheets play a crucial role in the ongoing effort to develop more accurate and reliable Earth system models for probabilistic sea-level projections. These simulations often require a massive amount of memory and computation from large supercomputing clusters to provide sufficient accuracy and resolution; therefore, it has become essential to ensure performance on these platforms. Many of today’s supercomputers contain a diverse set of computing architectures and require specific programming interfaces in order to obtain optimal efficiency. In an effort to avoid architecture-specific programming and maintain productivity across platforms, the ice-sheet modeling code known as MPAS-Albany Land Ice (MALI) uses high-level abstractions to integrate Trilinos libraries and the Kokkos programming model for performance portable code across a variety of different architectures. In this article, we analyze the performance portable features of MALI via a performance analysis on current CPU-based and GPU-based supercomputers. The analysis highlights not only the performance portable improvements made in finite element assembly and multigrid preconditioning within MALI with speedups between 1.26 and 1.82x across CPU and GPU architectures but also identifies the need to further improve performance in software coupling and preconditioning on GPUs. We perform a weak scalability study and show that simulations on GPU-based machines perform 1.24–1.92x faster when utilizing the GPUs. The best performance is found in finite element assembly, which achieved a speedup of up to 8.65x and a weak scaling efficiency of 82.6% with GPUs. We additionally describe an automated performance testing framework developed for this code base using a changepoint detection method. The framework is used to make actionable decisions about performance within MALI. We provide several concrete examples of scenarios in which the framework has identified performance regressions, improvements, and algorithm differences over the course of 2 years of development.
In this paper, we introduce PYSIMFRAC, an open-source python library for generating 3-D synthetic fracture realizations, integrating with fluid simulators, and performing analysis. PYSIMFRAC allows the user to specify one of three fracture generation techniques (Box, Gaussian, or Spectral) and perform statistical analysis including the autocorrelation, moments, and probability density functions of the fracture surfaces and aperture. This analysis and accessibility of a python library allows the user to create realistic fracture realizations and vary properties of interest. In addition, PYSIMFRAC includes integration examples to two different pore-scale simulators and the discrete fracture network simulator, dfnWorks. The capabilities developed in this work provides opportunity for quick and smooth adoption and implementation by the wider scientific community for accurate characterization of fluid transport in geologic media. We present PYSIMFRAC along with integration examples and discuss the ability to extend PYSIMFRAC from a single complex fracture to complex fracture networks.
Exchanging halo data is a common task in modern scientific computing applications and efficient handling of this operation is critical for the performance of the overall simulation. Tausch is a novel header-only library that provides a simple API for efficiently handling these types of data movements. Tausch supports both simple CPU-only systems, but also more complex heterogeneous systems with both CPUs and GPUs. It currently supports both OpenCL and CUDA for communicating with GPGPU devices, and allows for communication between GPGPUs and CPUs. The API allows for drop-in replacement in existing codes and can be used for the communication layer in new codes. This paper provides an overview of the approach taken in Tausch, and a performance analysis that demonstrates expected and achieved performance. Here, we highlight the ease of use and performance with three applications: First Tausch is compared to the halo exchange framework from two Mantevo applications, HPCCG and miniFE, and then it is used to replace a legacy halo exchange library in the flexible multigrid solver framework Cedar.
We examine a multi-modal approach to educating and training users of an advanced computing technology testbed at the Institute for Advanced Computational Science at Stony Brook University. Ookami provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes, this being the same processor technology powering the Japanese Fugaku supercomputer, the fastest computer in the world since June 2020. However, achieving high-performance on this Arm-based, leadership computing technology requires that users be familiar with details of computer architecture, performance analysis and modeling, and high-performance programming models that are commonly omitted in introductory programming courses. Indeed, regardless of their seniority, many of the testbed users are surprisingly unfamiliar with basic concepts such as vectorization, pipelining, latency/bandwidth, roofline models, computing energy/power, threads, and non-uniform memory access. These same concepts also pervade mainstream x86 technologies, so this is of widespread concern. Due to the national/global nature of our user community that is also very diverse in both discipline and experience, the inability to offer formal classes, and our experience that most people do not tend to read online documentation or training materials in sufficient depth, we have consciously employed multiple approaches that heavily emphasize (online) personal interactions and transfer of skills. Online documentation has been organized around best-practices and FAQs; twice-weekly hackathons and office hours via Zoom enable deep dives by both the team and the user community with multiple broad benefits; a Slack channel provides both real time and archived answers and discussions; and workshops, training and webinars target community needs as they arise. Furthermore, the perspective that these tools are being used in an educational setting rather than just for project communication makes them more effective and contributes to community success.
This poster summarizes a screening-level analysis performed by NETL's Strategic Systems Analysis and Engineering (SSAE) Directorate in support of achieving DOE's Hydrogen Shot Goal. Included are objectives, results, and primary conclusion of the analysis.
Protein Engineering is a highly evolved field of engineering aimed at developing proteins for specific industrial, medical, and research applications. Here, we present a practical teaching course to demonstrate fundamental techniques used to express, purify and analyze a recombinant protein produced in Escherichia coli—the enhanced green fluorescent protein (eGFP). The methodologies used for eGFP production were introduced sequentially over six laboratory sessions and included (i) bacterial growth, (ii) sonication (for cell lysis), (iii) affinity chromatography and dialysis (for eGFP purification), (iv) bicinchoninic acid (BCA) and fluorometry assays for total protein and eGFP quantification, respectively, and (v) sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) for qualitative analysis. All groups were able to isolate the eGFP from the cell lysate with purity levels up to 72%. Additionally, a mass balance analysis performed by the students showed that eGFP yields up to 46% were achieved at the end of the purification process following the adopted procedures. A sensitivity analysis was performed to pinpoint the most critical steps of the downstream processing.
Large language models (LLMs) have arisen rapidly to the center stage of artificial intelligence as the foundation models applicable to many downstream learning tasks. However, how to effectively build, train, and serve such models for many high-stake and first-principle-based scientific use cases are both of great interests and of great challenges. Moreover, pre-training LLMs with billions or even trillions of parameters can be prohibitively expensive not just for academic institutions, but also for well-funded industrial and government labs. Furthermore, the energy cost and the environmental impact of developing LLMs must be kept in mind. Here, in this work, we conduct a first-of-its-kind performance analysis to understand the time and energy cost of pre-training LLMs on the Department of Energy (DOE)’s leadership-class supercomputers. Employing state-of-the-art distributed training techniques, we evaluate the computational performance of various parallelization approaches at scale for a range of model sizes, and establish a projection model for the cost of full training. Our findings provide baseline results, best practices, and heuristics for pre-training such large models that should be valuable to HPC community at large. We also offer insights and optimization strategies for using the first exascale computing system, Frontier, to train models of the size of GPT-3 and beyond.
Electro-fuels can be produced from concentrated sources of carbon dioxide and hydrogen using electricity generated from renewable sources; this process enables energy storage at high volumetric energy density. Among the electro-fuels options, FT (Fischer-Tropsch) fuel is attractive for heavy-duty trucks and non-road transportation applications. This study conducts a techno-economic analysis of FT liquid fuel production from H 2 and CO 2 using a detailed performance analysis. Minimum fuel selling price is estimated for a broad range of H 2 and CO 2 prices and for a range of potential CO 2 credits. The analysis indicates that H 2 price has the largest impact on the minimum selling price of FT fuel. FT fuel production with a CO 2 price of $17.3/metric ton requires an H 2 price of $0.8/kg to be cost-competitive with the pre-tax petroleum diesel price of $3.1/gal in 2050 (before the application of any CO 2 credits). When the H 2 price is $2.0/kg from central water electrolysis (2020 target), the minimum selling price of the FT fuel is $5.4–5.9/gal. A sensitivity analysis shows that future system optimization of FT fuel production could focus on improving the H 2 and CO 2 recycle contributions and FT fuel conversion ratio. The analysis results can be combined with various upstream systems for H 2 and CO 2 production.
Developed at Lawrence Livermore National Laboratory (LLNL), ROSE is an open source compiler infrastructure to build source-to-source program transformation and analysis tools for large-scale C (C89 to C23), C++ (C++98 to C++23), UPC, Fortran (Fortran4, 66, 77, 95, 2003), OpenMP, Java, Python, and Binary applications. ROSE users range from experienced compiler researchers to library and tool developers who may have minimal compiler experience. ROSE is particularly well suited for building custom tools for static analysis, program optimization, arbitrary program transformation, domain-specific optimizations, complex loop optimizations, performance analysis, and cyber-security. ROSE is: A library (and set of associated tools) to quickly and easily apply compiler techniques to one's code in order to improve application performance and developer productivity. A research and development compiler infrastructure for for writing custom source-to-source translators to perform source code transformations, analysis, and optimizations. Is
SAND2024-00946O Python codes used in "Trade-offs in the latent representation of microstructure evolution," a manuscript accepted for publication in Acta Materialia, are part of a repository. The code was developed to perform analysis of microstructure evolution. The repository consists of two main directories: models, which train and test models such as autoencoders and diffusion maps, and analysis, which analyzes microstructures. Source code is used to perform dimensionality reduction of microstructure data for analysis of its evolution in time. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
Energy systems resilience is becoming increasingly important as the frequency of major grid outages increases. In this work, we present a methodology to optimize a behind-tlie-meter distributed energy resource system to sustain a site's critical loads during a pre-defined outage period. With the fixed system design, we then propose an outage simulation approach to estimate the resilience potential of the DER system to sustain loads beyond the fixed outage period - a yearlong resilience performance analysis. We apply statistical analysis to assess the system's resilience performance over a broader parametric problem space on an hourly, monthly, and yearly basis. We demonstrate the impact of the pre-defined outage period on the resilience performance through a case study. Results show that the probability of surviving a random outage of a given duration changes from 20% to 95% when the outage is modeled for a weekday instead of a weekend for the given load-profile.
The DeepLynx DAG repository will contain several Airflow DAGs (Directed Acyclic Graphs) which will be used in the context of DeepLynx's deployed Apache Airflow instance. These DAGs will be used for multiple data management tasks for DeepLynx data, including but not limited to: - bringing data from various sources and tools into DeepLynx - managing sequential data workflows, such as running Python scripts on data to perform analysis and returning the results to DeepLynx - performing any necessary transformation or pre-processing on data coming into DeepLynx from external sources or out of DeepLynx to go to external applications
We present a search for low-mass narrow qq̅ resonances. This search uses data from LHC pp collisions at a center of mass of 13 TeV in Run 2, and corresponds to an integrated luminosity of 137 fb^{-1}, currently using 10\% of data. Utilizing full Run 2 data allows the use of a lower photon pT threshold trigger than a previous analysis performed with only 2016 data, allowing this analysis to be more sensitive to resonances in the low mass region. We require an initial state photon recoiling against the narrow resonance, leading to the resonance having a high transverse momentum. The high pT decay products of the resonance collimate and are reconstructed as a single large jet with an internal two-pronged substructure. A two-pronged dijet score based on the ParticleNet tagger is used to select jets with two-pronged substructure. The background is estimated via a data-driven method using a transfer factor between the distributions which fail and pass the two-pronged substructure requirement. The new physics signal is searched for as a narrow peak excess above the Standard Model backgrounds in the jet mass spectrum.