Monday, August 31, 2026

Module 1.2 - Data Quality: Standards

In this lab assignment 1.2 Data Quality – Standards, we measured the quality of two road networks: the streets of Albuquerque, derived from a shapefile provided by the city of Albuquerque, and the Street Map USA data sourced from the TeleAtlas product. The objective was to evaluate the positional accuracy using metrics established by the National Standards for Spatial Data Accuracy (NSSDA). 

The application of the NSSDA standard involves seven distinct steps. 

The first step is to determine whether the assessment apply to horizontal accuracy, vertical accuracy, or both. For this lab, our focus is solely on determining the horizontal accuracy.

The second step requires selecting a study area and identifying a set of test points within that area. These test points correspond to street intersections, ensuring that we select good intersections and include a minimum of 20 street intersection points. The distribution of these points must be such that, when the study area is divided into four quadrants, each quadrant contains 20% of the test points. I selected 24 test points, ensuring that each quadrant included 6 test points. Furthermore, the distance between these points should exceed one-tenth of the length of the study area’s diagonal. Given that the diagonal length of my study area was 10 miles, I ensured that my test points were spaced at least 1 mile apart. The test points were chosen from both Albuquerque Street data and Street Map USA data correspond to the same street intersection. Both layers of test points are populated with identical point IDs that match to the same intersection. 


Screenshot of study area and sampling locations

The third step involves selecting a higher accuracy dataset for comparison with the data collected from the test area. Initially, we created a new layer designated as Reference, utilizing high-resolution aerial photography to digitize all points that correspond to the true locations of the test points. It was imperative to populate the ID field with the same test point number across all three datasets: Albuquerque Street, Street Map USA, and the Reference true point layer.

To determine the coordinate values for each test point for both the test data sets and the reference data set, I ran Calculate Geometry Attributes tool from Data Management on all three layers to add two new fields for the X and Y coordinates for the test points. At this point, we have three datasets of the same street intersections and have a common field “Point ID”.

Our next step is to assess the accuracy of both the Albuquerque Street layer and the Street Map USA layer. To achieve this, I employed the join tool in ArcGIS Pro to first join the Albuquerque Street layer with the Reference layer based on the common field, Points ID. Subsequently, I applied the join tool to join the Street Map USA test points layer with the Reference layer. Finally, I exported both layers into an Excel spreadsheet.

Using NSSDA accuracy statistic worksheet to calculate the positional accuracy statistic. 

For this step I added five new fields to the excel table they are:

DIFF IN X = the difference between the X coordinate from test point and reference true point

SQUARED DIFF IN X (1) = This is the square root of the x difference 

DIFF IN Y = the difference between the Y coordinate from test point and reference true point

SQUARED DIFF IN Y (2) = This is the square root of the Y difference 

(1) + (2) = SQUARED DIFF IN X + SQUARED DIFF IN Y (2)

From field (1) + (2), three key values can be calculated: The sum, the average and the root mean square error (RMSE).

The sum is the total of the squared differences between the coordinate values of the test data set and those of the Reference data set. The average is obtained by dividing the sum by the number of test points. The root mean square error (RMSE) is the square root of the average. 

The NSSDA statistic is then calculated by multiplying the RMSE by a factor that signifies the standard error of the mean at a 95 percent confidence level which is 1.7308 for the horizontal accuracy. Finally, NSSDA accuracy statement is prepared in a standardized report form.

NSSDA accuracy statement for Albuquerque Street Data:

Horizontal Positional Accuracy: Tested 23.25 feet horizontal accuracy at 95% confidence level

NSSDA accuracy statement for Street Map USA Data:

Horizontal Positional Accuracy: Tested 325.69 feet horizontal accuracy at 95% confidence level





Sunday, August 23, 2026

Module 1.1 - Calculating Metrics for Spatial Data Quality

Part A: Determining accuracy and precision

Precision indicates the degree to which the collected points are in proximity to one another; it is considered high when the points are closely grouped. Conversely, accuracy assesses how near the points are to the actual reference location. Average accuracy is deemed high when the mean of the points is situated close to the true location. 

In this laboratory exercise, we initially computed the estimated horizontal precision for a set of points gathered using a handheld GPS device. By applying the 68th percentile, we determined that the horizontal precision was 4.47 meters. This suggests that the points are clustered and are closed to each other. 

Subsequently, we assessed the horizontal distance from the average waypoint to the reference point, which represents the true location of the mapped point, in order to determine the horizontal accuracy, which was found to be 3.24 meters. This suggest that the points are near the true location. This minimal discrepancy indicates a high level of accuracy. 


 

Part B: Root-mean-square error (RMSE) and cumulative distribution function (CDF)

Error metrics for the GPS positions:

Minimum = 0.14
Maximum = 6.95
Mean = 2.67
Median = 2.45
RMSE = 3.06
68th Percentile = 3.18
90th Percentile = 4.67
95th Percentile = 5.69




In the first section of part B of the lab, we used metrics to calculate the RMSE (Root Mean Square Error), and also calculated the Mean, Average, minimum, maximum, 68th percentile, 90th percentile and 95th percentile values for error_xy. The error_xy is the distance error from a given point to the benchmark point. In the second section of part B, we created a cumulative distribution function (CDF) graph showing the complete distribution instead of the selected metrics. To accomplish that we used the error_xy data, sorting by smallest to largest values and adding a new column which represent the percentage from 0 to 100. We created a CDF graph showing the entire probability distribution of the error_xy. We can say that the CDF is a distributional description of data, while metrics are summary statistics derived from data.

The cumulative distribution function (CDF) provides significantly greater insight than metrics. The slope of the CDF indicates variations in density. In this analysis, the curve rises steeply until approximately 68%, where the majority of points are situated near the true location, while the flatter sections signify lower density with fewer points present. Additionally, the CDF offers a visual representation of cumulative distribution, enabling the comparison of multiple distributions. In summary, the CDF enables a more detailed interpretation than merely depending on a mean or median obtained from metrics.