Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “small files”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Implementation of the Windowed Multipole Method in Shift

The windowed multipole (WMP) method has been implemented in the Shift Monte Carlo (MC) radiation transport code with support for both CPU and GPU execution. With this method, small WMP data libraries (~100 MB) can be used to accurately Doppler broaden cross sections to arbitrary temperatures “on the fly” during an MC simulation. This approach yields significant memory savings relative to traditional methods, making it ideal for high-fidelity analysis such as coupled multiphysics simulations. This document provides the exact forms of the WMP equations used by Shift, as well as a detailed description of the structure of WMP HDF5 data files provided by the Massachusetts Institute of Technology (MIT). The Shift implementation has been validated against the OpenMC radiation transport code, with excellent agreement demonstrated for 70 nuclides across an operative range of temperatures. CPU and GPU performance testing using a small module reactor (SMR) problem demonstrated that this method decreases the neutron tracking rate by a factor of ~2 on the Summitdev machine. A new set of WMP data being developed in-house will employ novel methods to improve tracking rates.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The waterSHED Model: User Guide

The ideal design and operation of small hydropower plants is a complex optimization problem with economic, social, and environmental objectives. The waterSHED (Water Allocation Tool Enabling Rapid Small Hydropower Environmental Design) model is a user-friendly tool that allows hydropower stakeholders to model the trade-offs among these objectives using the Standard Modular Hydropower (SMH) framework. The SMH framework employs modular technologies that can be represented as blackbox objects and combined within a river to create a hydropower facility. For a given site, the waterSHED model aims to determine which modules should be placed in a facility and how those modules should be operated. This user guide describes how to use the graphical user interface and related functionalities. This document also summarizes the background research and mathematical formulations that are explained indepth in the accompanying doctoral dissertation. This model is an early step toward a new hydropower design process that employs standardization and modularity to reduce costs, development timelines, and challenges regarding social and environmental mitigation measures for low-head, small hydropower development. The waterSHED model is a Python application that will require the ability to download a GitHub repository, import the necessary packages, and run a set of Python script files using an integrated development environment. The script produces a graphical user interface to coordinate inputs, simulate operation, and visualize results, so no coding experience is needed once the script is running. Additionally, the waterSHED Workbook is a Microsoft Excel file that works with the Python script to facilitate data entry

13 HYDRO ENERGY↗

Analyses of the Constant Rate Discharge Test for 299-W15-225 (2009) and Ringold Formation Unit A Well Development Tests for 699-43-67B (2012) and 699-45-67B (2013)

This Environmental Calculation File (ECF) documents analyses of hydraulic tests performed in the Ringold Formation member of Wooded Island unit A (Rwia) and Ringold Formation member of Wooded Island unit E (Rwie) in the 200 West Area. Two small-scale tests in the Rwia performed in 2012 and 2013 and one large-scale hydraulic test in the Rwie performed in 2009 were analyzed. Heterogeneity in the Rwie was also examined.

54 ENVIRONMENTAL SCIENCES↗

Design and implementation of dynamic I/O control scheme for large scale distributed file systems

In this paper, we have analyzed the input/output (I/O) activities of Cori, which is a high-performance computing system at the National Energy Research Scientific Computing Center at Lawrence Berkeley National Laboratory. Our analysis results indicate that most users do not adjust storage configurations but rather use the default settings. In addition, owing to the interference from many applications running simultaneously, the performance varies based on the system status. To configure file systems autonomously in complex environments, we developed DCA-IO, a dynamic distributed file system configuration adjustment algorithm that utilizes the system log information to adjust storage configurations automatically. Our scheme aims to improve the application performance and avoid interference from other applications without user intervention. Moreover, DCA-IO uses the existing system logs and does not require code modifications, an additional library, or user intervention. To demonstrate the effectiveness of DCA-IO, we performed experiments using I/O kernels of real applications in both an isolated small-sized Lustre environment and Cori. Our experimental results shows that our scheme can improve the performance of HPC applications by up to 263% with the default Lustre configuration.

97 MATHEMATICS AND COMPUTING↗

TUMME: Tsinghua University Minnesota Master Equation program

We report that TUMME is a program for assembling and solving master equations for gas-phase chemical kinetics based on chemically significant eigenmodes. TUMME has interfaces to the Gaussian, Polyrate, and/or MSTor output files that allow the master equation code to obtain the microcanonical flux coefficients needed for the coefficient matrix of the master equation. The flux coefficients for reactions with barriers can be calculated by multi-structural variational transition state theory with small-curvature tunneling (MS-VTST/SCT) or by simpler approximations to this such as conventional transition state theory without tunneling (also called RRKM theory). The flux coefficients for barrierless reactions are provided by a hard-sphere model. TUMME is written in double precision with Python 3; quadruple and octuple precision are also available for some subtasks in C++. The Python code can run in serial or parallel (MP or MPI), and the C++ code can run on a single processor or on multiple processors with OpenMP.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Automated Production of Optimization-Based Control Logics for Dynamic Facade Systems, with Experimental Application to Two-Zone External Venetian Blinds

The primary goal of this research is to devise a system that produces controllers for complex fenestration systems that perform nearly as well as Model Predictive Control but at a level of cost and implementation complexity that rivals simple heuristic controls. To this end, a cloud-based automated controller production system has been set up for a motorized external Venetian blind device, with a simple web interface that can be used by non-experts. The computation cost per controller is in the range of a few dollars, and the control logic is simple enough to be implemented on small and cheap distributed controllers. The web interface allows the user to specify some details of their particular building and window configuration, including orientation, latitude, interior geometries, and lighting and HVAC system parameters. Upon submittal, a cloud-based system configures the necessary files and commands, and then runs thousands of optimizations with them. Once the calculations are finished, the system produces a lookup table and interpolation-based controller scripts that can be used on a simple and cheap distributed controller. This paper describes the underlying models and optimization processes. It also describes the resulting control logics for two cases tested at Lawrence Berkeley National Laboratory’s Advanced Windows Testbed Facility: illuminance maximization subject to glare constraints; and lighting + HVAC energy minimization. The performance of the model-based controllers produced by the automated web-based system are compared to a heuristic ‘block beam’ controller in physical experiments at the Testbed. The experimental results are supplemented by simulation experiments with the same configuration as the Testbed. The results show the illuminance maximizing controller significantly outperforms the heuristic controller in terms of glare avoidance, and also outperforms it in terms of hours of daylight autonomy. The energy minimizing controller also outperforms the heuristic controller. This paper also discusses how the web-based system may be extended to consider other configurations, such as electrochromic windows and thermally massive HVAC systems. Potential roles for this type of system within the building design and construction industry are discussed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Crystal Diffraction Prediction and Partiality estimation using Gaussian basis functions

A small dataset with good data quality, useful as a benchmark. There exists a two-fold indexing ambiguity, so when processing with CrystFEL you will need to do: ambigator -o disambiguated.stream -y m-3 -w m-3m -j 72 ambiguous.stream Please check the README file for more information about the dataset. ----- Begin unit cell ----- CrystFEL unit cell file version 1.0 lattice_type = cubic centering = I a = 103.40 A b = 103.40 A c = 103.40 A al = 90.00 deg be = 90.00 deg ga = 90.00 deg ; Please note: this is the target unit cell. ; The actual unit cells produced by indexing depend on many other factors. ----- End unit cell -----

Cydia pomonella granulovirus hull protein↗

cldera-tools

SAND2025-03848O CLDERA-Tools is a small library for performing online calculation of quantity of interest derived from variables in the host application. The library stores pointers to the arrays of variables of the host model and uses them at every timestep to compute desired quantity of interests (prescribed via yaml input files). Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Watkins II, Jerry↗

Barge Site - Avian Radar System / Derived Data

This is a combined data set of 67,410 bird/bat tracks from an avian radar system deployed on a research barge (MERLIN True3D, DeTect, Panama City, Florida, USA) and concurrent wind measurements from two scanning lidars (WindCube v2.1, Vaisala, Vantaa, Finland, and Halo XR+, Halo Photonics, Lannion, France). The research barge (16.5 m x 61 m) was deployed as part of the Wind Forecast Improvement Project (WFIP-3) off the northeast coast of the United States south of Massachusetts (40.9 deg N, 70.79 deg W). This data set comprises 5 weeks of data between August 27th 2024 and September 27th 2024. Radar data were provided by DeTect and Lidar data were accessed through the Wind Data Hub (wfip3/barg.WINDPROF.z01.a0) The data have been filtered and sorted into two size groups ("big" and "small") based on a clustering approach. See Snortland, A., Clerc, J., Hein, C., & Cotter, E. (2025). Wind as Driver of Bird and Bat Abundance, Flight Direction, Altitude, and Speed on the North Atlantic Shelf. arXiv preprint arXiv:2511.14983 for complete details. Data are provided in 2 files: "Birds" and "Birds_hourly" Birds: This file contains information about each of the 67,410 flying animal tracks detected by the radar during the data collection period, including parameters measured by the radar and wind information interpolated from the lidar wind measurements. We note that the raw radar dataset contained 301,618 tracks; tracks in this processed dataset were filtered based on the requirements described in Snortland et al. (2025). Birds_hourly: This file contains timeseries of the number of tracks detected per hour over the course of the data collection period, including wind conditions and sun position for each hour. These data were used for generalized additive modeling in Snortland et al. (2025).

17 WIND ENERGY↗

Data for Rod et al., "Alternating salt and freshwater floods of coastal soils impact soil structure, hydraulic properties, and oxygen dynamics"

This dataset includes laboratory experiment data on soil structure, hydraulic properties, and oxygen dynamics associated with Rod et al. 2026 https://doi.org/10.1002/vzj2.70073. There are six data files from a lab-based flood simulation of either freshwater (FW) or alternating brackish saltwater (SW) and FW using soil cores from a coastal forest at the Smithsonian Environmental Research Center. For soil information please see the Location section of the metadata. Files include: CO2, surface chemistry, water retention, dissolved oxygen, and soil specific surface area. Each file is in CSV format and can be opened/read with any plain text tabular file reader (Microsoft Excel, R, etc.). Purpose of Experiment: To investigate how hydrologic intensification affects soil structure and oxygen dynamics, we conducted a series of laboratory-based flood simulations. After three SW-FW floods (6 floods total) there were significant changes in pore size distribution, significant redistribution of colloids, and the A-horizon became sodic. We concluded that a small number of SW flooding events can induce a measurable change in soil physical properties that directly impacts the biogeochemical dynamics.

54 ENVIRONMENTAL SCIENCES↗

2021 Smoky Mountains Conference Data Challenge Synthetic-to-Real Domain Adaptation for Autonomous Driving Dataset

The dataset is comprised of both real and synthetic images from a vehicle's forward-facing camera. Each camera image is accompanied by a corresponding pixel-level semantic segmentation image (all files are .png files). In total, the dataset contains 5600 images in the training/validation set and 1400 images in the testing set. The training dataset contains mostly synthetic RGB images collected with a wide range of weather and lighting conditions using the CARLA simulator [1]. In addition, the training data also includes a small pre-selected subset of data from the Cityscapes training dataset – which is comprised of RGB-segmentation image pairs from driving scenarios in various European cities [2]. The testing data is split into three sets. The first set contains synthetic CARLA images with weather/lighting conditions that were not present in the training set. The second set is a subset of the Cityscapes testing dataset. Finally, the third set is an unknown testing set which will not be revealed to the participants until after the submission deadline. [1] Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, October). CARLA: An open urban driving simulator. In Conference on robot learning (pp. 1-16). PMLR. [2] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... and Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213-3223).

99 GENERAL AND MISCELLANEOUS↗

SARS-CoV2 Docking Dataset

Description: Small-molecule conformations and docking scores for 1.4 billion molecules docked against 6 protein targets from SARS-CoV2: MPro 5R84, MPro 6WQF, NSP15 6WLC, PLPro 7JIR, Spike 6M0J, and a hand-optimized model of the RNA-dependent RNA polymerase. Docking was carried out using the Autodock-GPU program performing 20 independent structure minimizations per dock - saving 3 results per molecule. Scores reported include the Autodock free energy estimate as well as RF3 and VS-DUD-E v2 machine-learned rescoring models. Protein structure files and maps in the format input to Autodock-GPU are included. Literature Ref: Supercomputer-Based Ensemble Docking Drug Discovery Pipeline with Application to Covid-19, J. Chem. Inf. Model. 2020, 60(12): 5832–5852.

36 MATERIALS SCIENCE↗

Nuclear Science for the Manhattan Project and Comparison to Today’s ENDF Data

Nuclear physics advances in the United States and Britain from 1939 to 1945 are described. The Manhattan Project’s work led to an explosion in our knowledge of nuclear science. A conference in April 1943 at Los Alamos provided a simple formula used to compute critical masses and laid out the research program needed to determine the key nuclear constants. In short order, four university accelerators were disassembled and reassembled at Los Alamos, and methods were established to make measurements on extremely small samples owing to the initial lack of availability of enriched 235U and plutonium. I trace the program that measured fission cross sections, fission-emitted neutron multiplicities and their energy spectra, and transport cross sections, comparing the measurements with our best understanding today as embodied in the Evaluated Nuclear Data File ENDF/B-VIII.0. The large nuclear data uncertainties at the beginning of the project, which often exceeded 25% to 50%, were reduced by 1945 often to less than 5% to 10%. Uranium-235 and plutonium-239 fission cross-section assessments in the fast mega-electron-volt range were reduced following more accurate measurements, and the neutron multiplicity $\overline{v}$ increased. By a lucky coincidence of canceling errors, the initial critical mass estimates were close to the final estimated masses. Some images from historical documents from our Los Alamos archives are shown. Many of the original measurements from these early years have not previously been widely available. Through this work, these data have now been archived in the international experimental nuclear reaction data library (EXFOR) in a collaboration with the International Atomic Energy Agency and Brookhaven National Laboratory.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Code Description for "Brief Communication: Monitoring snow depth using small, cheap, and easy-to-deploy ground surface temperature sensors"

Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We train a random forest machine learning model to predict snow depth from variability in ground surface temperature. To our knowledge, this is the first time that small ground surface temperature sensors have been used to estimate snow depth. The model performs well at sites where the model was trained and at pan-arctic evaluation sites (RMSE <= 0.15 m). Small temperature sensors are cheap and easy-to-deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring to an extent previously infeasible. The model is flexible and can be applied to datasets retroactively to retrieve snow depth estimates at additional sites. This code package includes a *.joblib file of the trained random forest model and a *.ipynb file showing how to clean input data, train the random forest model, and apply the model.

Bachand, Claire↗

Forensic Analysis of SOHO Router Binaries

Small Office/Home Office (SOHO) routers are used by millions of consumers across the United States, and are commensurately vulnerable. Forensic analysis of SOHO router firmware helps to understand and mitigate those vulnerabilities. This poster focused particularly on analysis of BusyBox executables, a software suite that provides several Unix utilities in a single file. Three main tools were used to analyze the binaries. BinWalk was used to extract the files, but also to build entropy graphs, extract Linux kernel images, and identify CPU architectures; WiiBin processed the binaries to find endianness, architecture, the percent compressed/encrypted, and compiler data; and @DisCo, a machine learning tool used to determine function similarity in disassembled binaries, analyzed similarities and determined versions of extracted BusyBox files from each router. These tools found that venders from all five routers utilized the same version of the BusyBox software across different firmware updates, demonstrating the importance of constant firmware scrutiny to protect against security vulnerabilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine learning snow depth predictions at sites in Alaska, Norway, Siberia, Colorado and New Mexico

Temporally continuous snow depth estimates are vital for understanding changing snow patterns in the Arctic and impacts on permafrost. We trained random forest machine learning models to predict snow depth from temperature data recorded at or just below the ground surface. Training data was collected at the Teller 27 Watershed and Kougarok 64 Hillslope during the 2021 - 2022 water year on the Seward Peninsula, Alaska using distributed temperature profiling (DTP) systems. We then applied this model to other sites where ground surface or shallow soil temperature data was available for at least one water year (see Related Datasets). Many of these temperature measurements were collocated with snow depth observations. Ground surface temperature (i.e. snow-ground interface temperature) is easy to measure using small, cheap and easy-to-deploy temperature sensors such as iButtons and TinyTags, and such measurements have previously been used to calculate a variety of snow metrics (e.g. snow onset date). However, this is the first study to estimate snow depth directly from ground surface temperature data. The present dataset contains one *.csv file which includes machine learning snow depth predictions at sites in Alaska, Norway, Siberia, Colorado, and New Mexico and one *.kml file including the locations of sites with snow depth predictions. No training data predictions are included in the *.csv file. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy’s Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy’s Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

SAIL-Net Raw and Post Corrected POPS Data Fall 2021 - Summer 2023

SAIL-Net is a DOE funded project in the East River Watershed near Crested Butte, Colorado with the goal of advancing our understanding of aerosol-cloud interactions in complex, mountainous regions. Through the deployment of a network of six low cost microphysics nodes in Fall 2021 in the same domain at the SAIL campaign, SAIL-Net provides data on aerosol size distributions, cloud condensation nuclei (CCN), and ice nucleation particles (INP). This network enables the investigation of small-scale variations in complex terrain. Two datasets are provided - one containing raw data and the other containing post-corrected data. The raw dataset provides the raw data recorded from the POPS which were deployed at each of the six sites. These data are organized by site and broken down into daily data files. The six site names used here are: “gothic”, “irwin”, “cbtop”, “cbmid”, “pumphouse”, and “snodgrass”. These data are not cleaned or post-corrected, but some flags have been added. The data are reported at 1 second time resolution. The post-corrected dataset provides the post-corrected and cleaned data recorded from the POPS which were deployed at each of the six sites. This data are also organized by site (same as those found in the raw data) and broken down into daily data files. Unlike the raw POPS data, these data have already been cleaned to remove what we believe are bad values. These data should be ready to use with no cleaning. For a full description of the cleaning and post-correction process, see the readme.

54 ENVIRONMENTAL SCIENCES↗