Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as abrupt time series data shifts. These shifts, which are usually unintentional, may be caused by software or hardware configuration changes on a PV system, and do not reflect an actual change in overall system performance. Locating these shifts and segmenting the associated time series aids in more accurate future PV analysis. In this research, an offline changepoint detection (CPD) algorithm that automatically detects these abrupt data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window for the labeled data sets. To promote further research in this space, we are releasing the labeled data shift sets on U.S. Department of Energy's (DOE) DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package.

changepoint detection↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series: Preprint

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as abrupt time series data shifts. These shifts, which are usually unintentional, may be caused by software or hardware configuration changes on a PV system, and do not reflect an actual change in overall system performance. Locating these shifts and segmenting the associated time series aids in more accurate future PV analysis. In this research, an offline changepoint detection (CPD) algorithm that automatically detects these abrupt data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window for the labeled data sets. To promote further research in this space, we are releasing the labeled data shift sets on U.S. Department of Energy's (DOE) DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package.

changepoint detection↗

Innovating the next generation of commercial smart building software

Nearly 30% of commercial building energy use is wasted due to equipment faults and HVAC controls problems. The result is increased emissions, compromised comfort and productivity, and less reliable coordination of building power needs with a clean grid. The energy impact alone represents $17 billion in potential savings. Today’s smart building software provides a robust solution to address these operational deficiencies. Energy management and information systems (EMIS) are saving up to 9% on average, with two-year paybacks. They are being incorporated into energy management processes, commissioning services, and utility programs. As effective as they are, two barriers prevent even deeper benefits; limited personnel to fix problems once they are identified, and the expense and time to manually implement changes in control systems. In partnership with the research community, the EMIS industry is developing new capabilities to overcome these barriers. Moving beyond siloed products for either fault detection and diagnostics, or optimal control, these new capabilities empower users to not only automatically identify faults, but also to push corrective action, and control improvements to their buildings. In this paper, several areas for enhancements are documented: ‘one-time’ correction of faults such as setpoints, schedules, and economizer lockouts; short-term active testing for automated proportional integral derivative (PID) loop tuning and functional testing; and continuous supervisory control for demand flexibility and year-round efficiency. Results are presented from a pair of partner implementations out of a dozen providers integrating these enhancements into their products, including field tests from across the country, and insights into operator acceptance and integration into operations and maintenance practices.

Casillas, Armando↗

Breaking the barrier of human-annotated training data for machine learning-aided plant research using aerial imagery

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

59 BASIC BIOLOGICAL SCIENCES↗

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Use of Machine Learning on PMU Data for Transmission System Fault Analysis

Synchrophasor technology has been used for monitoring, control, and protection of bulk power system for over 10 years. Deployment of phasor measurement units (PMUs) in the USA power system has surpassed 3000 units installed in the transmission substations as stand-alone intelligent electronic devices (IEDs) or as a software add-on to other devices such as digital protective relays (DPRs) or digital fault recorders (DFRs). By now, thousands of terabytes of PMU data may have been captured and stored by various transmission system operators (TSOs) and independent system operators (ISOs). This creates an opportunity to deploy advanced machine learning (ML) techniques to detect and classify faults recorded by PMUs automatically to be used by the system operators for rapid, critical decision-making when manual analysis of the past or unfolding events is not feasible. In this paper we offer a brief background on how the automated fault analysis may be done using DPR and/or DFR data, and compare some of the legacy approaches to the new ML approaches in the context of the system-wide PMU recordings. We then offer insights from developing practical ML solutions that have been applied on field recordings captured by close to 450 PMUs from all three US interconnections (Western, Eastern and ERCOT) over two years (2016-2017). We identify and illustrate ML challenges we addressed: inaccurate data, data with scarce and temporally imprecise fault labels, data recorded by PMUs sparsely located at substations resulting in the fault records taken afar from the ends of the faulted lines, data containing only positive sequence values, and data taken at different voltage levels. We then illustrate the ML model results for fault analysis under different application scenarios. The novelty of this study is not only in the design, implementation, and performance analysis of the ML algorithms, but also in the use of advanced fault modelling and simulation approaches to improve the training results when developing supervised ML models for fault detection and classification. Extensive simulations of faults were conducted on a 14-bus power system to create a training dataset with over 1400 accurately labelled faults. This dataset was applied to enhance the accuracy of fault detection and classification of machine learning-based models trained with small number of labelled faults in large datasets recorded in the grid interconnections ranging from 5,000 to 70,000 buses.

Synchrophasors, Machine Learning, Fault Analysis, ↗

Machine Learning-Based Identification of the Interface Regions for Coupling Local and Nonlocal Models

Local-nonlocal coupling approaches provide a means to combine the computational efficiency of local models and the accuracy of nonlocal models. However, the coupling process can be challenging, requiring expertise to identify the interface between local and nonlocal regions. Here, this study introduces a machine learning-based approach to automatically detect the regions in which the local and nonlocal models should be used. The method uses loading functions evaluated at grid points to decide the model selection at those points. Training of the networks is based on datasets provided by classes of loading functions for which reference coupling configurations are computed using accurate coupled solutions, where accuracy is measured in terms of the relative error between the solution to the coupling approach and the solution to the nonlocal model. We study two approaches that vary in data structure. The first, the full-domain input data approach, uses the entire load vector and outputs a complete label vector, performing a global classification. The second, a window-based approach, processes loads into windows and addresses the problem as a node-wise classification where each window's central point is classified individually. The classification problems are solved via deep learning algorithms based on convolutional neural networks. The performance of these approaches is studied on one-dimensional numerical examples using F1-scores and accuracy metrics. Notably, the windowing approach achieves an accuracy of 0.96 and an F1-score of 0.97, highlighting its potential to automate coupling processes effectively and enhance computational efficiency in material science applications.

97 MATHEMATICS AND COMPUTING↗

Data Driven Commercial Building Energy Code Compliance and Technology Inventory for New York City

Building Performance Standards (BPS) are gaining national traction. A BPS will require new processes in the design, construction, and operation of buildings that take the occupants into account and enable predictive analysis to ensure compliance with current and future GHG emissions caps. In New York City, most buildings over 25,000 square feet will be regulated by a BPS starting in 2024, regardless of whether it is new construction permitted under current energy codes or an existing building. This research is one of the first to begin the evaluation of a long-term series of building policies in the context of an open data ecosystem, in cooperation with city agencies. Existing building policies enacted in NYC have ranged from building energy benchmarking and labeling to energy audits to the regulation of GHG emission in buildings. Through the development of a dataset related to building technologies and energy consumption, this project can help to evaluate if meaningful conclusions can be drawn for the data that has been largely self-reported in compliance with city regulations. This project will also provide lessons learned from a deep dive into these types of datasets to provide best practices for municipalities or states seeking to embark on policies like those enacted in NYC. In addition, a Building Automation System (BAS) Stretch Standard of Care (SSOC) for owners, designers, and building operators will enable the measurement and predictive analysis of energy consumption and GHG emissions at the plant, system, or component level, in anticipation of regulated GHG limits on buildings based on energy use. The SSOC is expected to be suitable for use on a national level. The primary feature of an SSOC is a standardized format for a set of BAS points that can be used to control and to gather data from individual plants, systems, or components that are related to building energy consumption. This project examined how measurements compare to prescriptive or simulation-based energy code targets, finding little correlation between predictive 8760-hour energy modeling and actual energy consumption for a small sample (n=27) of buildings constructed after 2015. Other analysis found that, while large multifamily housing (MFH) buildings showed a general trend similar to predicted reductions in energy use from the implementation of model commercial energy codes, this trend was not evident in the office, K-12 school, and hotel use groups in NYC. No upward or downward trends in energy consumption were found when buildings were grouped by size. Energy audit data were analyzed and it appears that there is bias by audit company on measures recommended to clients. Further research should be performed to cross-analyze this with other attributes, such as building size, vintage, and number of stories. Analysis found that for 281 buildings that were permitted and completed after 2015 and had submitted benchmarking data in 2022, between 81% and 96% (by use group) were found to be in compliance with the 2024 to 2029 NYC BPS emission caps, and between 55% and 89% were in compliance with the 2030-2034 caps. This work is beneficial to the public in helping policymakers and building stakeholders better understand the wide-ranging implications of a BPS.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS↗

Model Agnostic Bayesian Framework for Online Anomaly/Event Detection in PMU Data

Phasor measurement units (PMU) are integral to the modernization and automation plan of the electric power industry. A PMU data signature contains system-level events (e.g., faults, generation/load change, etc.) and any measurement/device-related errors. Therefore, the reliable and resilient operation of power systems is equivalent to the quality of the PMU data and the situation awareness provided by its data signature. Despite recent progress, current state-of-the-art methods are not fool-proof and have certain limitations tracing an error/abnormality to sensor sub-components and grid systems. This is because of technical challenges imposed by the scarcity of the labeled information, loss of data quality, and non-stationarity of data. In this paper, we consider the online PMU data stream as an output of a stochastic process and pose the anomaly/event detection as a changepoint detection problem dealing with detecting parameter changes in the underlying stochastic processes. The proposed model-agnostic framework relies on: (a) feature extraction utilizing the minimum volume enclosing ellipsoids (MVEE) method from raw PMU observations and (b) a Bayesian framework of changepoint detection. The validity of the proposed methodology is discussed through numerical experiments on real-world utility-scale PMU data.

Hossain, Ramij Raja↗

Open Data and Deep Semantic Segmentation for Automated Extraction of Building Footprints

Advances in machine learning and computer vision, combined with increased access to unstructured data (e.g., images and text), have created an opportunity for automated extraction of building characteristics, cost-effectively, and at scale. These characteristics are relevant to a variety of urban and energy applications, yet are time consuming and costly to acquire with today’s manual methods. Several recent research studies have shown that in comparison to more traditional methods that are based on features engineering approach, an end-to-end learning approach based on deep learning algorithms significantly improved the accuracy of automatic building footprint extraction from remote sensing images. However, these studies used limited benchmark datasets that have been carefully curated and labeled. How the accuracy of these deep learning-based approach holds when using less curated training data has not received enough attention. The aim of this work is to leverage the openly available data to automatically generate a larger training dataset with more variability in term of regions and type of cities, which can be used to build more accurate deep learning models. In contrast to most benchmark datasets, the gathered data have not been manually curated. Thus, the training dataset is not perfectly clean in terms of remote sensing images exactly matching the ground truth building’s foot-print. A workflow that includes data pre-processing, deep learning semantic segmentation modeling, and results post-processing is introduced and applied to a dataset that include remote sensing images from 15 cities and five counties from various region of the USA, which include 8,607,677 buildings. The accuracy of the proposed approach was measured on an out of sample testing dataset corresponding to 364,000 buildings from three USA cities. The results favorably compared to those obtained from Microsoft’s recently released US building footprint dataset.

97 MATHEMATICS AND COMPUTING↗

End-to-End Pipeline for Trigger Detection on Hit and Track Graphs

There has been a surge of interest in applying deep learning in particle and nuclear physics to replace labor-intensive offline data analysis with automated online machine learning tasks. This paper details a novel AI-enabled triggering solution for physics experiments in Relativistic Heavy Ion Collider and future Electron-Ion Collider. The triggering system consists of a comprehensive end-to-end pipeline based on Graph Neural Networks that classifies trigger events versus background events, makes online decisions to retain signal data, and enables efficient data acquisition. Here, the triggering system first starts with the coordinates of pixel hits lit up by passing particles in the detector, applies three stages of event processing (hits clustering, track reconstruction, and trigger detection), and labels all processed events with the binary tag of trigger versus background events. By switching among different objective functions, we train the Graph Neural Networks in the pipeline to solve multiple tasks: the edge-level track reconstruction problem, the edge-level track adjacency matrix prediction, and the graph-level trigger detection problem. We propose a novel method to treat the events as track-graphs instead of hit-graphs. This method focuses on intertrack relations and is driven by underlying physics processing. As a result, it attains a solid performance (around 72% accuracy) for trigger detection and outperforms the baseline method using hit-graphs by 2% higher accuracy.

97 MATHEMATICS AND COMPUTING↗

Single‐Domain Multiferroic Array‐Addressable Terfenol‐D (SMArT) Micromagnets for Programmable Single‐Cell Capture and Release

Abstract Programming magnetic fields with microscale control can enable automation at the scale of single cells ≈10 µm. Most magnetic materials provide a consistent magnetic field over time but the direction or field strength at the microscale is not easily modulated. However, magnetostrictive materials, when coupled with ferroelectric material (i.e., strain‐mediated multiferroics), can undergo magnetization reorientation due to voltage‐induced strain, promising refined control of magnetization at the micrometer‐scale. This work demonstrates the largest single‐domain microstructures (20 µm) of Terfenol‐D (Tb 0.3 Dy 0.7 Fe 1.92 ), a material that has the highest magnetostrictive strain of any known soft magnetoelastic material. These Terfenol‐D microstructures enable controlled localization of magnetic beads with sub‐micrometer precision. Magnetically labeled cells are captured by the field gradients generated from the single‐domain microstructures without an external magnetic field. The magnetic state on these microstructures is switched through voltage‐induced strain, as a result of the strain‐mediated converse magnetoelectric effect, to release individual cells using a multiferroic approach. These electronically addressable micromagnets pave the way for parallelized multiferroics‐based single‐cell sorting under digital control for biotechnology applications.

Khojah, Reem↗

Artificial intelligence for materials research at extremes

Abstract Materials development is slow and expensive, taking decades from inception to fielding. For materials research at extremes, the situation is even more demanding, as the desired property combinations such as strength and oxidation resistance can have complex interactions. Here, we explore the role of AI and autonomous experimentation (AE) in the process of understanding and developing materials for extreme and coupled environments. AI is important in understanding materials under extremes due to the highly demanding and unique cases these environments represent. Materials are pushed to their limits in ways that, for example, equilibrium phase diagrams cannot describe. Often, multiple physical phenomena compete to determine the material response. Further, validation is often difficult or impossible. AI can help bridge these gaps, providing heuristic but valuable links between materials properties and performance under extreme conditions. We explore the potential advantages of AE along with decision strategies. In particular, we consider the problem of deciding between low-fidelity, inexpensive experiments and high-fidelity, expensive experiments. The cost of experiments is described in terms of the speed and throughput of automated experiments, contrasted with the human resources needed to execute manual experiments. We also consider the cost and benefits of modeling and simulation to further materials understanding, along with characterization of materials under extreme environments in the AE loop. Graphical abstract AI sequential decision-making methods for materials research: Active learning, which focuses on exploration by sampling uncertain regions, Bayesian and bandit optimization as well as reinforcement learning (RL), which trades off exploration of uncertain regions with exploitation of optimum function value. Bayesian and bandit optimization focus on finding the optimal value of the function at each step or cumulatively over the entire steps, respectively, whereas RL considers cumulative value of the labeling function, where the latter can change depending on the state of the system (blue, orange, or green).

36 MATERIALS SCIENCE↗

FY23 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), and in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or all of: height, color, and grayscale values as functions of position in a plane projection) to detect the presence of surface corrosion and cracking after being trained on similar data with the features to be detected labeled.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A self-supervised robotic system for autonomous contact-based spatial mapping of semiconductor properties

Integrating robotically driven contact-based material characterization techniques into self-driving laboratories can enhance measurement quality, reliability, and throughput. While deep learning models support robust autonomy, current methods lack reliable pixel-precision positioning and require extensive labeled data. To overcome these challenges, we propose an approach for building self-supervised autonomy into contact-based robotic systems that teach the robot to follow domain expert measurement principles at high throughputs. We demonstrate the performance of this approach by autonomously driving a 4-DOF robotic probe for 24 hours to characterize semiconductor photoconductivity at 3025 uniquely predicted poses across a gradient of drop-casted perovskite film compositions, achieving throughputs of more than 125 measurements per hour. Spatially mapping photoconductivity onto each drop-casted film reveals compositional trends and regions of inhomogeneity, valuable for identifying manufacturing defects. With this self-supervised neural network–driven robotic system, we enable high-precision and reliable automation of contact-based characterization techniques at high throughputs, thereby allowing measurement of previously inaccessible yet important semiconductor properties for self-driving laboratories.

Science & Technology - Other Topics↗

Total Dissolved Nitrogen and Ammonia Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for total dissolved nitrogen (TDN) and ammonia concentrations for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. TDN was analyzed using a Shimadzu Total Nitrogen Module (TNM-1) combined with the TOC-VCSH analyzer (Shimadzu Corporation, Japan). TNM-1 is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. Ammonia was determined using a Lachat's QuikChem 8500 Series 2 Flow Injection Analysis System (LACHAT Instruments, QuckChem 8500 series 2, Automated Ion Analyzer, Loveland, Colorado). When ammonia in water samples is heated (60 degrees C) with salicylate and hypochlorite in an alkaline phosphate buffer, an emerald green color is produced which is proportional to the ammonia concentration. The color is intensified by the addition of nitroprusside. Ethylenediaminetetraacetic acid (EDTA) is added to the buffer to prevent the interference of metal ions (Ca, Mg, and Fe etc.). Ammonia-N is then determined by LACHAT flow injection and a colorimetric assay at an absorbance wavelength 660 nm. (Reference: LACHAT Instruments: QuickChem Method 90-107-06-3-A, Determination of Ammonia by Flow Injection Analysis (High Throughput, Salicylate Method/DCIC) (Multi Matrix method). Written by Lynn Egan (Application group), February 08, 2011.) All files are labeled by location and variable, and data reported are the mean values upon replicate measurements. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process as detailed in the methods. This data package contains (1) a zip file (tdn_ammonia_data_2015-2025.zip) containing a total of 299 files: 298 data files of ammonia and TDN data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the determination of Method Detection Limits (MDLs) for TDN data, which has been updated in 2026-08; and (5) PDF and docx files for the detemination of Method Detection Limits (MDLs) for Ammonia and the Interferences by LACHAT Flow Injection Analysis. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 105 locations containing TDN and Ammonia-N data. Update 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses and Determination of Method Detection Limit for Ammonia and the Interferences by LACHAT Flow Injection Analysis documents, which can be accessed as PDFs or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated total dissolved nitrogen and ammonia data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Units were listed incorrectly, but have been fixed to reflect correct units (ug/L). File level metadata (flmd) and data dictionary (dd) files were updated to reflect the updated versions of these files. Available data was added up until 2022-06-01. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-27. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for TDN were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for TDN data were added to this dataset.

54 ENVIRONMENTAL SCIENCES↗