Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Automatic Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SolarAPP+ Performance Review (2023 Data)

The Solar Automated Permit Processing Plus (SolarAPP+) platform is an online portal to facilitate and expedite rooftop solar photovoltaic (PV) and battery storage permitting processes. SolarAPP+ allows PV contractors to upload system specifications, have that information automatically reviewed for code compliance, and receive instant approval for code-compliant systems, reducing authority having jurisdiction (AHJ) staff time needed for review. SolarAPP+ also provides inspection checklists to verify installation practices and adherence to approved designs. SolarAPP+ is available to AHJs at no cost. This report is part of an ongoing series of reviews of SolarAPP+ performance. Consistent with previous performance reviews, we summarize SolarAPP+ adoption trends to date and compare various metrics for PV systems permitted through SolarAPP+ versus systems permitted through traditional AHJ permitting processes. As of the end of 2023, the National Renewable Energy Laboratory (NREL) had contacted over 1,700 AHJs with significant solar permitting volume regarding SolarAPP+. Of those, 793 AHJs had expressed interest in the platform as of the end of 2023. 161 AHJs had begun piloting the platform and 97 of these had publicly launched the platform by the end of 2023. In 2023, 668 installers submitted 18,906 permits through the SolarAPP+ platform, including 4,834 permits submitted as part of a solar plus storage program. SolarAPP+ permits accounted for around 43% of all permits issued in participating AHJs. We compare permitting timelines through SolarAPP+ to traditional AHJ permitting processes to assess the platform's performance. Consistent with previous SolarAPP+ performance reviews, we find that permitting timelines are significantly shorter for SolarAPP+ projects. Based on median timelines, a typical SolarAPP+ project is permitted and inspected 14.5 business days sooner than traditional projects. We estimate that automatic SolarAPP+ permitting saved around 7,200 hours of AHJ staff time in 2023. Finally, we estimate that SolarAPP+ eliminated over 150,000 business days in permitting-related delays in 2023.

14 SOLAR ENERGY

SolarAPP+ Performance Review (2024 Data)

The Solar Automated Permit Processing Plus (SolarAPP+) platform is an online portal to facilitate and expedite rooftop solar photovoltaic (PV) and battery storage permitting processes. SolarAPP+ allows PV contractors to upload system specifications, have that information automatically reviewed for code compliance, and receive instant approval for code-compliant systems, reducing authority having jurisdiction (AHJ) staff time needed for review. SolarAPP+ also provides inspection checklists to verify installation practices and adherence to approved designs. This report is part of an ongoing series of reviews of SolarAPP+ performance. Consistent with previous performance reviews, we summarize SolarAPP+ adoption trends to date and compare various metrics for PV systems permitted through SolarAPP+ versus systems permitted through traditional AHJ permitting processes. As of the end of 2024, 799 AHJs had expressed interest in the platform, with 264 fully adopting (215) or piloting (49) the platform. In 2024, 861 installers submitted 37,393 permits through the SolarAPP+ platform, including 27,375 permits for PV+storage systems. SolarAPP+ permits accounted for around 43% of all permits issued in all participating AHJs, and more than 60% of all permits in several participating AHJs. We compare permitting timelines through SolarAPP+ to traditional AHJ permitting processes to assess the platform's performance. Consistent with previous SolarAPP+ performance reviews, we find that permitting timelines are significantly shorter for SolarAPP+ projects. Based on median timelines, a typical SolarAPP+ project is permitted and inspected 12 business days sooner than traditional projects. We estimate that automatic SolarAPP+ permitting saved around 18,400 hours of AHJ staff time in 2024. Finally, we estimate that SolarAPP+ eliminated over 100,000 business days in permitting-related delays in 2024.

14 SOLAR ENERGY

Impact of Color Space and Color Resolution on Vehicle Recognition Models

In this study, we analyze both linear and nonlinear color mappings by training on versions of a curated dataset collected in a controlled campus environment. We experiment with color space and color resolution to assess model performance in vehicle recognition tasks. Color encodings can be designed in principle to highlight certain vehicle characteristics or compensate for lighting differences when assessing potential matches to previously encountered objects. The dataset used in this work includes imagery gathered under diverse environmental conditions, including daytime and nighttime lighting. Experimental results inform expectations for possible improvements with automatic color space selection through feature learning. Moreover, we find there is only a gradual decrease in model performance with degraded color resolution, which suggests the need for simplified data collection and processing. By focusing on the most critical features, we could see improved model generalization and robustness, as the model becomes less prone to overfitting to noise or irrelevant details in the data. Such a reduction in resolution will lower computational complexity, leading to quicker training and inference times.

47 OTHER INSTRUMENTATION

Event Log / Raw Data

The WFIP3 event log is a curated record spanning 578 days of meteorological phenomena and field observations that complements the campaign’s high-frequency measurements. The log combines manually documented daily weather discussions with automatically derived indicators of key atmospheric processes, providing standardized, publicly available context to support model evaluation, forecast verification, and case-study selection for offshore boundary-layer research.

17 WIND ENERGY

NeuNorm

NeuNorm is a scipp-based Python library for neutron imaging normalization and time-of-flight (TOF) data processing at Oak Ridge National Laboratory imaging facilities (MARS at HFIR and VENUS at SNS). NeuNorm 2.0 is a complete, scipp-based rewrite of the original NeuNorm normalization library, adding HDF5 output, automatic uncertainty propagation, and TOF/event-mode processing.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations

Data and scripts associated with a manuscript modeling microbial regulation of priming effects

This data package is associated with the publication “Modeling Microbial Regulatory Feedback in Organic Matter Decomposition Identifies Copiotrophic Traits as Key Drivers of Positive Priming” published as a preprint on BioRXiv by Ahamed et al. (2026); https://doi.org/10.1101/2024.08.11.607483. The package contains MATLAB scripts and saved simulation outputs used to implement a cybernetic model of microbial regulation during complex organic matter (OM) decomposition governing priming effects. It includes models of (i) single microbial functional groups (copiotrophic or oligotrophic degraders) and (ii) binary consortia composed of degraders and non-degraders with contrasting or common growth traits. Simulation results were generated using Monte Carlo analyses, with randomized key model parameters across a range of environmental mixing fractions of complex and labile OM. The dataset was created to provide a transparent and reusable computational framework for systematically exploring how microbial growth traits, metabolic regulation, and community composition influence OM decomposition dynamics and priming effects. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes the variable definitions. This package includes: (1) annotated MATLAB code implementing the system of ordinary differential equations and cybernetic control laws; (2) saved output files containing data (e.g., biomass, substrates, enzyme levels, priming metrics); and (3) scripts for processing saved outputs and regenerating figures. Specifically, the data package contains three main MATLAB scripts: runPrimingModel.m, runPlotData.m, and runPlotSuppFigS1.m, along with this readme and supporting documentation. Users should begin with runPrimingModel.m, which contains the annotated code implementing the system of ordinary differential equations and cybernetic control laws. This script runs the Monte Carlo simulations of microbial OM decomposition and allows users to modify microbial trait definitions, adjust parameter distributions, or define new community configurations. Simulation outputs are automatically saved as .mat files in the folder named SavedData, which stores all pre-generated results included in this package. The second script, runPlotData.m, reads files from the SavedData folder and processes them to regenerate the figures presented in the manuscript. The third script, runPlotSuppFigS1.m, specifically generates Figure S1 in the Supplementary Material of the manuscript. The package also includes the aforementioned files in non-proprietary .txt format. If users intend to use them, they should first save the files in their respective .m or .mat formats prior to execution in MATLAB.

Biomass concentration

XRF-XFS-XAS-Auto v1.0 - Beta release

This software allows to analyze XRF maps, XFS spectra and XAS spectra collected at the Advanced Light Source's Beamline 10.3.2. Features include: 1) XRF maps: - process XRF maps, all elemental maps are saved as bmp automatically and labeled with the incident energy used, the scale bar is also labeled and can be controlled. - XRF elemental correlation plots, save the correlation plots automatically - Extract single or multiple transects in XRF maps on one or several regions of interest, each transect profile is numbered and saved in a corresponding folder, along with the corresponding maps showing transect location. 2) XFS spectra - save in log10 scale the XFS spectra, either a single or multiple files all at once. The files are saved as .bmp. - XFS spectra are labeled according to tabulated fluorescence emission lines. 3) XAS spectra - allows to plot individual scalers in the raw data. - allows calibration of the spectra using an Io internal glitch present in all spectra and performing 1st derivative. - Least-square linear combination fitting of XANES or extended XANES spectra using a database of standards using 1, 2 or 3 components maximum. It also provides the 5 top combinations and provide the user for the possibility of saving the 2nd, 3rd, 4th and 5th best combinations in addition to the best one. The processed spectra (pre-edge background substracted, post-edge normalized), the fits and residuals are automatically saved. A table of the component, with fit% and SSN is provided and saved automatically as well.

Fakra, Sirine

Understanding Event Trajectories Across Massive Temporal Datasets with Word Embeddings and Visualization

In collaboration with researchers from Virginia Tech, Savannah River National Laboratory has continued development of a natural language processing pipeline to identify and extract events of interest from massive open data sources in the domain of worldwide state-sponsored civil nuclear energy. The foundation of the pipeline is built on compass aligned temporal word embedding models, whereby contextual shifts are automatically identified by comparing keyword embedding vectors across successive time windows. Within the approach, a contextual shift indicates the occurrence of a potential event of interest. However, in such a broad topical domain that captures events at a global scale, across various life cycle stages, and across numerous different technology types, a user that is monitoring events may have broad interests in capturing many different event types with varying degrees of signal. As such, the quantity of information that may be returned from an automated event extraction pipeline can be substantial, requiring manual effort to sift through the information to identify any relevant bits of information. Therefore, a more streamlined workflow that aids in directing a user toward specific information at different points in time is necessary. The workflow presented here has been developed with this concept in mind, built on top of the initial prototype event extraction pipeline, whereby a user can analyze temporal text-based data sources at multiple different contextual levels to isolate key points in time and key subdomains captured within a data corpus. Using multiple corpuses that consist of approximately 7 million Tweets and 7 million news articles, the team has extended compass aligned temporal word embedding models to establish an interconnected and hierarchical structure that relates known key words of interest to documents, local topics (i.e., within a time window), and global topics across the corpuses. All of this information is packaged into a visual analytics system that is linked to the information extraction pipeline and enables a user to identify contextual information that describes the evolution of a high dimensional embedding space across time to isolate changes of interest and explore associated events. This report demonstrates the use of these analytics and a means to fuse information across multiple datasets.

97 MATHEMATICS AND COMPUTING

A semi–automatic analytical methodology for characterizing the energy consumption of MRI systems using load duration curves

Background and purpose: Magnetic resonance imaging (MRI) scanners are a major contributor to greenhouse gas emissions from the healthcare sector, and efforts to improve energy efficiency and reduce energy consumption rely on quantification of the characteristics of energy consumption. The purpose of this work was to develop a semi-automatic analytical methodology for the characterization of the energy consumption of MRI systems using only the load duration curve (LDC). LDCs are a fundamental tool used across various fields to analyze and understand the behavior of loads over time. Methods: An electric current transformer sensor and data logger were installed on two 3T MRI scanners from two vendors, termed M1 (outpatient scanner) and M2 (inpatient/emergency scanner). Data was collected for 1 month (7/11/2023 to 8/11/2023). Active power was calculated, assuming a balanced three-phase system, using the average current measured across all three phases, a 480 V reference voltage for both machines, and vendor-provided power factors. An LDC was constructed for each system by sorting the active power values in descending order and computing the cumulative time (in units of percentage) for each data point. The first derivative of the LDC was then computed (LDC’), smoothed by convolution with a window function (sLDC’), and used to detect transitions between different system modes including (in descending power levels): scan, prepared-to-scan, idle, low-power, and off. The final, segmented LDC was used to measure time (% total time), total energy (kWh), and mean power (kW) for each system mode on both scanners. The method was validated by comparing mean power values, computed using the segmented 1-month LDC, for each nonproductive system mode (i.e., prepared-to-scan, idle, lower-power, and off) against power levels measured after a deliberate system shutdown was performed for each scanner (1 day worth of data). Results: The validation revealed differences in mean power values <1.4% for all nonproductive modes and both scanners. In the scan system mode, the mean power values ranged from 29.8 to 37.2 kW and the total energy consumed for 1 month ranged from 11 106 to 14 466 kWh depending on the scanner. Over the course of 1 month, the portion of time the scanners were in nonproductive modes ranged from 76% to 80% across scanners and the nonproductive energy consumption ranged from 8010 to 6722 kWh depending on the scanner. The M1 (outpatient) scanner consumed 99.9 and 183.9 kWh/day in idle mode for weekdays and weekends, respectively, because the scanner spent 23% more time proportionally in idle mode on the weekends. Conclusions: A semi-automatic method for quantifying energy consumption characteristics of MRI scanners was introduced and validated. This method is relatively simple to implement as it requires only power data from the scanners and avoids the technical challenges associated with extracting and processing scanner log files. Finally, the methodology enables quantitative evaluation of the power, time, and energy characteristics of MRI scanners in scan and nonproductive system modes, providing baseline data and the capability of identifying potential opportunities for enhancing the energy efficiency of MRI scanners.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

An iterative method to deblend AGN-Host contributions for Integral Field spectroscopic observations

ABSTRACT We present a new iterative deblending method to separate the host galaxy (HG) and their Active Galactic Nuclei (AGNs) emission with the use of Integral Field spectroscopic (IFS) data. The method decomposes the resolved HG emission from the unresolved AGN emission by modelling the two-dimensional surface brightness (SB) profile of the point-spread function (PSF) and the two-dimensional SB HG continuum simultaneously per each monochromatic slide. Our method does not require any prior information about the observed SB profile or a detailed fitting of the PSF, making it ideal for the automatic analysis of large galaxy samples. In this work, we test the quality of our method, its advantages, and its disadvantages. We test our method by using a set of IFS mock data cubes to quantify the reliability of our deblending process and further compare our method with the qdblend3d analysis tool. Furthermore, we applied our method to three data cubes selected from the MaNGA survey according to the dominance of either its HG or its AGN. We show that our deblending method is capable of disengaging the bright, non-resolved AGN emission from the HG continuum and its narrow emission lines. However, the decoupling depends on how well the IFS spatially resolves the PSF, and on the relative flux intensity of the HG-AGN. Therefore, the method is ideal for disentangling the bright-flux contribution from AGN-dominated spectra.

Ibarra-Medel, H. (ORCID:0000000297906313)

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul [National Renewable Energy Lab. (NREL

Radiation image reconstruction and uncertainty quantification using a Gaussian process prior

We propose a complete framework for Bayesian image reconstruction and uncertainty quantification based on a Gaussian process prior (GPP) to overcome limitations of maximum likelihood expectation maximization (ML-EM) image reconstruction algorithm. The prior distribution is constructed with a zero-mean Gaussian process (GP) with a choice of a covariance function, and a link function is used to map the Gaussian process to an image. Unlike many other maximum a posteriori approaches, our method offers highly interpretable hyperparamters that are selected automatically with the empirical Bayes method. Furthermore, the GP covariance function can be modified to incorporate a priori structural priors, enabling multi-modality imaging or contextual data fusion. Lastly, we illustrate that our approach lends itself to Bayesian uncertainty quantification techniques, such as the preconditioned Crank–Nicolson method and the Laplace approximation. The proposed framework is general and can be employed in most radiation image reconstruction problems, and we demonstrate it with simulated free-moving single detector radiation source imaging scenarios. We compare the reconstruction results from GPP and ML-EM, and show that the proposed method can significantly improve the image quality over ML-EM, all the while providing greater understanding of the source distribution via the uncertainty quantification capability. Furthermore, significant improvement of the image quality by incorporating a structural prior is illustrated.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Subject-specific modeling framework for particle deposition using computational fluid dynamics

Quantifying particle deposition and dose in the respiratory tract requires a physiologically realistic representation and reproducible computational workflows. However, existing modeling frameworks, such as the International Commission on Radiological Protection (ICRP) compartmental models and the Multiple Path Particle Dosimetry (MPPD) tool, lack detailed deposition profiles and subject-specific capabilities. The combination of advances in computer vision algorithms applied to the respiratory tract and Computational Fluid and Particle Dynamics (CFPD) allows high-fidelity simulations of particle behavior in anatomically accurate geometries derived from individual CT scans. The segmentation, preprocessing, and file preparation task for a CFPD simulation was often time-consuming, and no prior studies to-date have yet presented a fully automated framework. This work presents a fully automated workflow to obtain individualized particle deposition profiles in the human respiratory tract. The pipeline starts with segmenting upper and lower airway geometries using morphological and deep learning-based methods, generating three-dimensional (3D) models from CT imaging data. Next, a series of algorithms are presented to quality check and prepare the 3D geometry for a CFD or CFPD simulation. The preprocessing step includes correcting geometric artifacts, enforcing a physically consistent mesh, and automatically identifying and capping multiple outlets, which is required for CFD/CFPD simulations. These processed models are then input into open-source (OpenFOAM) or commercial (StarCCM+) CFD solvers, where flow and transient particle transport equations — including turbulence and particle–wall interactions are solved under realistic breathing conditions. Finally, the resulting particle deposition profiles can be integrated with Monte Carlo radiation transport codes and state-of-the-art computational phantoms to assess organ-specific absorbed doses in scenarios of radioactive aerosol inhalation. The presented work streamlines respiratory tract segmentation, preprocessing for CFD/CFPD simulations, and integration with dose assessment workflows, reducing manual intervention and improving access to high-fidelity, subject-specific modeling. The high precision in predicted particle deposition and dose distributions can improve personalized treatment strategies in respiratory medicine and refine dose estimates for radiation protection.

AI

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie

Distribution Substation Planning Toolkit (dsp-toolkit) v1.0

The Distribution Substation Planning Toolkit (DSP Toolkit) is a software suite designed to streamline the planning and optimization of distribution substations. This toolkit offers a comprehensive set of tools and APIs for data curation, short-term electric load forecasting, and weather-sensitive load adjustment, making it an essential resource for utility companies, engineers, and researchers. Features • Data Preprocessing and Curation: Efficiently manage and preprocess large datasets to ensure high-quality input for analysis. • Short-Term Load Forecasting: Utilize data-driven models to predict short-term electric loads accurately. • Weather-Sensitive Modeling: Automatically adjust load forecasts based on weather data to predict future peak demands more precisely. Uses The DSP Toolkit is ideal for planning and optimizing distribution substations, providing a user-friendly interface and comprehensive documentation. It is suitable for both novice and experienced users, facilitating efficient and accurate planning processes. Advantages • Efficiency: Automates complex planning tasks, reducing manual effort and minimizing errors. • Scalability: Handles large datasets and complex models, making it suitable for large-scale projects. • Community and Support: Open-source with active community contributions, ensuring continuous improvement and support. • Extensibility: Easily extendable with custom modules and plugins, allowing users to tailor the toolkit to their specific needs. The DSP Toolkit stands out by offering a robust, flexible, and user-friendly solution for distribution substation planning. Public Abstract

Li, Han [Lawrence Berkeley National Laboratory (LB

SymbolFit: Automatic Parametric Modeling with Symbolic Regression

We introduce SymbolFit (API: https://github.com/hftsoi/symbolfit), a framework that automates parametric modeling by using symbolic regression to perform a machine-search for functions that fit the data while simultaneously providing uncertainty estimates in a single run. Traditionally, constructing a parametric model to accurately describe binned data has been a manual and iterative process, requiring an adequate functional form to be determined before the fit can be performed. The main challenge arises when the appropriate functional forms cannot be derived from first principles, especially when there is no underlying true closed-form function for the distribution. In this work, we develop a framework that automates and streamlines the process by utilizing symbolic regression, a machine learning technique that explores a vast space of candidate functions without requiring a predefined functional form because the functional form itself is treated as a trainable parameter, making the process far more efficient and effortless than traditional regression methods. We demonstrate the framework in high-energy physics experiments at the CERN Large Hadron Collider (LHC) using five real proton-proton collision datasets from new physics searches, including background modeling in resonance searches for high-mass dijet, trijet, paired-dijet, diphoton, and dimuon events. We show that our framework can flexibly and efficiently generate a wide range of candidate functions that fit a nontrivial distribution well using a simple fit configuration that varies only by random seed, and that the same fit configuration, which defines a vast function space, can also be applied to distributions of different shapes, whereas achieving a comparable result with traditional methods would have required extensive manual effort.

Tsoi, Ho Fung [Univ. of Pennsylvania, Philadelphia

Reconstruction framework advancements to support streaming for the ePIC detector at the EIC

The ePIC collaboration adopted the JANA2 framework to manage its reconstruction algorithms. This framework has since evolved substantially in response to ePIC’s needs. There have been three main design drivers: integrating cleanly with the Podio-based data models and other layers of the key4hep stack, enabling external configuration of existing components, and supporting timeframe splitting for streaming readout. The result is a unified component model featuring a new declarative interface for specifying inputs, outputs, parameters, services, and resources. This interface enables the user to instantiate, configure, and wire components via an external file. One critical new addition to the component model is a hierarchical decomposition of data boundaries into levels such as Run, Timeframe, PhysicsEvent, and Subevent. Two new component abstractions, Folder and Unfolder, are introduced in order to traverse this hierarchy, e.g. by splitting or merging. The pre-existing components can now operate at different event levels, and JANA2 will automatically construct the corresponding parallel processing topology. This means that a user may write an algorithm once, and configure it at runtime to operate on timeframes or on physics events. Overall, these changes mean that the user requires less knowledge about the framework internals, obtains greater flexibility with configuration, and gains the ability to reuse the existing abstractions in new streaming contexts.

Brei, Nathan [Thomas Jefferson National Accelerato