Leveraging high-performance data transfer to offload data management tasks to SmartNICs
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Small Scale WEC Performance Modeling Data is performance data from downscaled models of common WEC devices and their calculated performance outputs. This data is used by the Small WEC interactive modeling tool hosted by PRIMRE. The devices include a point absorber, a two-body point absorber (RM3), an oscillating surge device (OSWEC), and an attenuator type device (McCabe Wave Pump). One of the primary use cases for this work is to give an easy way to compare power output for a variety of WECs and model sizes.
A techno-economic analysis is underway examining the cost and performance of future large-scale photovoltaic (PV) plant components, including bifacial modules, tandem modules, increased plant voltage architectures, and module-level power electronics. Integration of these components into PV plant designs is compared with current PV technologies based on levelized cost of electricity (LCOE). Baseline models are developed and validated against recorded PV plant performance data. Expected cost and performance data of future PV technologies are incorporated into the baseline models. An evolutionary algorithm is utilized to optimize PV plant configuration, technology combination, and LCOE. Furthermore, this paper focuses on the bifacial module analysis.
Increases in data volumes are forcing high-energy and nuclear physics experiments to store more frequently accessed data on tape. Extracting the maximum performance from tape drives is critical to make this viable from a data availability and system cost standpoint. The nature of data ingest and retrieval in an experimental physics environment make achieving high access performance difficult given the inherent limitations of magnetic tape. Tailoring the layout of data on tape is one key to improving read performance. This paper highlights the work in progress to characterize ATLAS data ingested in the tape system, understand how data layout, i.e. file co-location on tape and file distribution over tapes, affect read performance and how optimal data layout might be achieved in a production environment.
Large-scale photovoltaic system performance analysis is being conducted within the US Department of Energy's-sponsored PV Fleet Performance Data Initiative. This collaboration with commercial PV system owners collects and evaluates PV field performance data, and provides reports on aggregated results. Drawing on over 2200 sites across the US and over 24,000 separate PV inverters we have collected in excess of 8.3 gigawatts (GW) of performance data, representing 6-7% of the entire US installed PV capacity. A mixture of utility-scale and large commercial systems are represented, averaging 4.1 megawatts (MW) in size and 5 years in age. Initial results show average system degradation rates at -0.75% / year, which is slightly higher than historically reported module-level values of -0.5%/year. We also found that the availability of systems averaged 97.7%, which is lower than the typical 99% uptime assumed by many project economic forecasts. Given these results, we compared monthly performance with expected production values, based on satellite weather data and a simple PVWatts performance model. We found that systems were performing within 10% of monthly expectation over 90% of the time, with a fleet average vs expected monthly value of 0.994. We also evaluated the impact of extreme weather events on system performance, and found a range of short-term and longer-term performance effects ranging from grid outage, system downtime, module damage and accelerated long-term degradation rate.
Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.
The Photovoltaic (PV) Performance Modeling Collaborative (PVPMC) organized a blind PV performance modeling intercomparison to allow PV modelers to blindly test their models and modeling ability against real system data. Measured weather and irradiance data were provided along with detailed descriptions of PV systems from two locations (Albuquerque, New Mexico, USA and Roskilde, Denmark). Participants were asked to simulate the plane-of-array irradiance, module temperature, and DC power output from six systems and submit their results to Sandia for processing. This dataset includes seven MS-Excel sheets with instructions, notes and all necessary data (weather, irradiance, temperature, power) used for the data analysis of the blind modeling comparison. The hourly data represent six different systems from Albuquerque, NM and Roskilde, Denmark over a period of one year. These data are useful for PV performance model validation studies.
Breakdown of the DOE Advanced Gas Reactor Fuel Development and Qualification Program. Including discussion topics on TRISO technology Status Circa 2000, New Production Reactor (NPR) Fuel Experience, Fuel Qualification, US DOE Advanced Gas Reactor (AGR) Fuel Development and Qualification Program, Initial AGR Program Reference HTGR Design, Fuel Qualification Approach, Fuel Performance Modeling, Fuel Fabrication, Selected AGR-1, AGR-2, and AGR-5/6/7 Fuel Property Means, AGR Program TRISO Fuel Key Performance Data, Irradiation Performance: Fission Gas R/B, Irradiation Testing Results, Kernel and Coating Behavior During Irradiation, Locating and Studying Failed Particles Greatly Improves Understanding of Fuel Performance, Fission Product Release from UCO Fuel Compacts: AGR-1 and AGR-2 Examples, HTGR Accident Safety Testing of TRISO Fuel, Evaluating Behavior During D-LOFC Accidents, Safety Test Results for US UCO Fuel, Particle Failure Evaluation, Fuel Performance Summary, Ongoing Work and Outstanding Data Needs, Core Oxidation, Industry Engagement, and Coated-Particle-Fueled Reactor Concepts and Fuel Designs.
The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.
The mapping of computational needs onto execution resources is, by and large, a manual task, and users are frequently guided simply by intuition and past experiences. We present a queueing theory based performance model for streaming data applications that takes steps towards a better understanding of resource mapping decisions, thereby assisting application developers to make good mapping choices. The performance model (and associated cost model) are agnostic to the specific properties of the compute resource and application, simply characterizing them by their achievable data throughput. We illustrate the model with a pair of applications, one chosen from the field of computational biology and the second is a classic machine learning problem.
Although model predictions of thermal energy storage (TES) performance have been explored in previous investigations, relevant test data that enable experimental validation of performance models have been limited. This is particularly true for high-performance TES designs that facilitate fast input and extraction of energy. In this paper, we present a summary of experimental tests of a high-performance TES unit using lithium nitrate trihydrate phase change material as a storage medium. Performance data are presented for complete dual-mode cycles consisting of extraction (melting) followed by charging (freezing). These tests simulate the cyclic operation of a TES unit for asynchronous cooling in a variety of applications. Finally, the model analysis is found to agree reasonably well, within 10%, with the experimental data except for conditions very near the initiation of freezing, a consequence of subcooling that is required to initiate solidification.
A techno-economic analysis is underway examining the cost and performance of future large-scale photovoltaic (PV) plant components, including bifacial modules, tandem modules, increased plant voltage architectures, and module-level power electronics. Integration of these components into PV plant designs is compared with current PV technologies based on levelized cost of electricity (LCOE). Baseline models are developed and validated against recorded PV plant performance data. Expected cost and performance data of future PV technologies are incorporated into the baseline models. An evolutionary algorithm is utilized to optimize PV plant configuration, technology combination, and LCOE. This paper focuses on the bifacial module analysis.
Tailored Fiber Placement (TFP) offers a novel approach to optimize fiber architecture for the fabrication of complex, structural parts not traditionally suitable for advanced composites. This technology not only offers new routes for weight reduction via metal substitution, it also offers cost reduction through minimization of material scrap and reduced labor. This reduction in component weight leads to increased fuel efficiency, and reduced production energy consumption, thereby, helping to achieve the stated IACMI technical goals. This technology leverages centuries of manufacturing development in support of the textile and embroidery industry. One major drawback to this technology is the lack of commercial or non- proprietary structural performance data and robust analytical tools used to optimize fiber architecture and predict performance. This project was structured to utilize common sub-element features to validate analytical performance tools, generate performance data, and gather cost and performance data on components of interest. This project was designed to give industry sponsors the confidence and ability to take full advantage of TFP to fabricate primary, highly loaded structure and integrate features such as metallic fasteners. The project focused principally on the use of high strength carbon fiber, such as T700, and the use of aerospace epoxy resin matrix to primarily support development of new composite applications in vehicle, aerospace, and industrial markets. This project applied previously developed analytical tools to predict the performance of TFP produced parts. This work focused on developing the pipeline to characterize material in order to accurately predict component performance when modifying the TFP print paths and stitch density. This focused on experimental characterization via standardized ASTM testing, alongside experimental testing of more representative service components by testing curved beam strength, beam shear performance, a large scale TFP lug, and ultimately designing a fully TFP clip bracket that reduced weight and cost compared to a traditional metallic component. The new knowledge gained from this program included: 1) development and demonstration of novel analytical tools applied to analysis of TFP preforms; 2) development and demonstration of a building block approach using coupons and sub-elements to optimize the design of a more complex component; 3) demonstration that optimized fiber orientation using TFP can exceed performance of conventional textile composite materials and can open new applications currently limited to metallic components; 4) Demonstration of performance and cost benefits of the TFP process as compared to metallic and conventional textile composites. Recommendations for follow-on work include development of design allowables to assess the impact of high temperature/moisture exposure or saturation during loading, tracking the impact of stitching needle wear on the performance of parts and ability to stitch thicker preforms, using TFP preforms as local reinforcement at areas of bearing or complex loading, and topology optimization of components by tow steering. The expertise developed during the course of this project can be leveraged to provide commercial engineering design and fabrication services using TFP. UDRI is in the process of formalizing their partnership with Spintech, who will serve as the commercialization partner for this technology and provide molding services and deliver finished components to the end user. UDRI will continue to produce the preforms until the economics allow Spintech to procure its own TFP equipment or lease UDRI equipment, at which point UDRI will step away from manufacture and serve as the engineering and design lead on product development.
Performing reliable Rietveld analysis on tens or hundreds of powder diffraction datasets from parametric or time-resolved experiments often poses a bottleneck in extracting meaningful results from the data. While automated analysis of data has recently been demonstrated, high temperature annealing studies, during which phase transformations occur and lattice parameters may change due to repartitioning of elements, are prime examples where automation by a simple phase identification from a database of room temperature structures or automation by sequential refinements is likely to fail. To enable reliable, efficient, automated Rietveld analysis, we present a Python package named Spotlight , building on established Rietveld packages such as MAUD, GSAS , or GSAS-II , which extends the refinement of best fit parameters to a global optimization using an ensemble of optimizers leveraging hierarchical parallel execution on high-performance computing clusters. Spotlight further enables the efficient design of refinement plans through the iterative automated machine-learning of a surrogate for the refinement on which the global optimizations are performed until results from the surrogate converge to the response surface data. We demonstrate Spotlight with the analysis of uranium molybdenum and Ti–6Al–4V datasets, as well as in two open-source tutorials analyzing aluminium oxide and lead sulphate.
Vehicle data consist of electric vehicle performance data collected directly from the vehicle during standard operations. Data were collected using onboard data loggers that were either installed by the project team or preinstalled by the original equipment manufacturer. Data recorded by the data loggers were made accessible via an online web portal or an application programming interface. Different data loggers were used (HEM, ViriCiti, and Geotab), and the method for each vehicle is defined in the vehicle attributes file. Some systems collected data on a “trip-level” basis, in which each row of a table represents a single trip (the period between a key-on and key-off event), whereas other data were collected on a per-day basis, in which each row represents a single day of operation. Data were collected over a range of data collection periods, depending on the project. Data have been anonymized by removing information or decreasing information resolution as necessary so that fleets are not identifiable. Due to the wide range of vehicle types represented and variation in data collection, data parameters and frequencies differ between vehicles and fleets The **Performance Data Daily/Trip Data Dictionaries** contain definitions for each available parameter associated with a vehicle’s operations, aggregated at either a daily or trip level. The parameters available will vary from vehicle to vehicle, but every possible parameter will be defined. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. The **Vehicle Data** tables contain the data from each vehicle’s operations, aggregated at either a daily or trip level, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.
Probabilistic Safety Assessment (PSA) of complex facilities is performed to arrive at the risk posed by them. PSA also accounts for the contribution of the human errors towards the overall risk through Human Reliability Analysis (HRA) in terms of Human Error Probability (HEP). Human operators are part of the system and do not work in isolation. Their performance is influenced by the context in which the actions are performed. As a result, quantification of HEP requires operator performance data under the given context. Some good sources of operator performance data are plant‘s operation data, simulator data and expert judgement. The plant operation data pertaining to HRA is generally sparse. In this situation, a full scope plant simulator provides a good alternative for operator performance data generation. Many of the currently practiced HRA methods have been developed by combining the empirical evidence with expert judgement and contain a lot of uncertainty in their estimates. Bayesian inference is suitable for updating the prior HRA estimates with the simulator evidence to obtain the posterior HEP. Here, posterior HEP has been calculated for postulated accident scenarios in advanced reactor (first of its kind) at design stage, using plant simulator.
This data encompasses performance data measured for flat-plate photovoltaic (PV) modules installed in Cocoa, Florida; Eugene, Oregon; and Golden, Colorado. The data include PV module current-voltage curves and associated meteorological data for approximately one-year periods. The data was acquired with the NREL Performance and Energy Rating Testbed (PERT) and the mobile Performance and Energy Rating Testbed (mPERT).
Performance characteristics of parallel particle advection algorithms can vary greatly based on workload.With this short paper, we build a new algorithm based on results from a previous bake-off study which evaluated the performance of four algorithms on a variety of workloads. Our algorithm, called HyLiPoD, is a ''meta-algorithm,'' i.e., it considers the desired workload to choose from existing algorithms to maximize performance. To demonstrate HyliPoD's benefit, we analyze results from 162 tests including concurrencies of up to 8192 cores, meshes as large as 34 billion cells, and particle counts as large as 300 million. Our findings demonstrate that HyLiPoD's adaptive approach allows it to match the best performance of existing algorithms across diverse workloads.