Thursday, October 3, 2024

Module 6- Special Topics- Scale Effect and Spatial Data Aggregation

Our final lab for Special Topics began with exploring the effects of scale on vector data and resolution on raster imagery. We began with a hydrographic vector data set scaled to 1:1,200, 1:24,000, and 1:100,000 and calculated the area and perimeter of the polygon features as well as the count and total length of the line features. The result of scaling the feature classes differently was that as the scale decreased, variables calculated for the water features increased. In addition to the length of the water features being calculated at a higher number, the difference in the scale was visually noticeable as some hydrographic features were only visible at smaller scale.

A close up view of my hydrographic line features with yellow being the largest scale, and blue being the smallest scale. As you can see, the smaller the scale the finer the features that visible.


After exploring scale we resampled a DEM file to compare 6 different resolutions: 1m, 2m, 5m, 10m, 20m, 90m. Raster resolution is correlated with pixel size, so smaller resolution numbers indicate smaller, more detailed pixels, while a raster with larger pixel sizes would be expected to have less detail. We calculated the slope of each raster and examined the results which showed that as the resolution increased, the average slope also decreased.

For the second half of the lab, we explored the Modified Area Unit Problem by using census data to explore how the results of our data analysis can differ based on how we group our data for analysis. In this situation we were analyzing census data to determine how race affected poverty status, and we separately analyzed data by block groups, zip codes, housing voting districts and counties. We saw that the data could vary significantly based on which of the 4 units of analysis were used to analyze it.

Lastly, we explored gerrymandering. Gerrymandering is the process of manipulating the boundaries of voting districts in order to favor a certain political party. Gerrymandered districts tend to have very irregular boundaries and one way to measure this is the Polsby-Popper Test, which uses and equation to measure how compact a district is on a scale of 0 (least compact) to 1 (most compact), with the assumption that the most compact districts are the least gerrymandered. Here is an examples of an extremely gerrymandered district as determined by this method.



Tuesday, September 24, 2024

Lab 5- Interpolation- Special Topics

Interpolation is a way of making predictions based on known values. We interpolated the surface water quality of Tampa Bay using the BOD Mg/L variable from multiple sampling locations.

Interpolation methods explored were
        - Spline (regular)
        - Spline (tension
        - IDW
        - Thiessen polygons

IDW uses the distance of each data point to neighboring values to weight its predictions, whereas the Spline interpolation techniques fits through a set of input points to make its estimation. Spline is not constrained by the limits of the provided data and can give values greater or less than the highest and lowest data points, while IDW is constrained to the highest and lowest data point. 

Thiessen interpolation creates a series of polygons with each data point associated with the central location of a polygon. Because of this the size of the polygons is variable and dependent on the distance from other sampling locations. Each polygon is given the value of the central point and there is not estimation or continuity to the data. This method was not ideal for measuring water quality.

In this exercise my regular spline interpolation ended up displaying values that were not mathematically possible, giving me large swaths of negative value areas. It is highly unlikely for water quality to be at 0, and impossible for it to be a negative number. The tension method was a more accurate and aside from a greater range of values, data was closer to that displayed by the IDW. I personally chose to map the data along standard water quality monitoring guidelines using the following table for analysis:

and when using these guidelines the spline (tension) method only slightly differed from the IDW with the majority of the data falling into the "very good" range for both data sets, but a slightly bigger circle area falling under "fair" in the spline tension map.

IDW: "Fair" data is mapped in yellow and "very good" data in green

Spline tension: "Fair" data is mapped in yellow and "very good" data in green

Finally, here is more thoroughly symbolized view of my IDW map with BOD estimates. Although I would not necessarily need to use these number for water quality estimation, they give a good visual of how the IDW interpolation method calculated the water quality with regard to the data points.


 



Sunday, September 15, 2024

Special Topics- Lab 4- TINS and DEMS

 This week we explored utilizing ArcGIS Pro to create elevation models. Additionally we converted a DEM to a TIN in order to create a ski run suitability map that weighted aspect, slope and elevation and presented ideal locations for a ski slope.

TINs and DEMs are terrain models used to analyze topographic features. TIN data uses a type of vector data to model a 3D surface by linking a series of triangles of different sizes based on elevation point data. These data points vary in density in proportion to the terrain the terrain. More complex terrain contains more data points, while simpler terrain contains fewer data points. A DEM consists of raster data consisting of a regular grid pattern that models elevation. 

We used our TIN and DEM models to depict the contour lines and slope of our data. Although DEM contour lines are smoother and more reminiscent of a typical map, TIN contour lines are more accurate due to the ability of the data to represent various surface features, notably peaks and ridges.

The first image displays my TIN while the second image contrasts my DEM layer of the same data.





The biggest difference I noticed between the two layers was that the DEM seemed to over estimate the uppermost regions of the data, adding contour lines where they were absent from the TIN data.



Wednesday, September 4, 2024

Special Topics- Data Quality Assessment- Module 3

 This week we were asked to assess the quality of two road network data sets, Street Centerlines and Tiger, by comparing the completeness of the data on a grid-by-grid basis. We then mapped the resulting data onto a choropleth map.

To run this assessment, I split the data so that each grid could be analyzed separately.  I then used spatial join and summary statistics tools to further process the data. I exported my statistics to Excel to more easily analyze it and found that the Street Centerlines had no data for one of my grids. In total the Street Centerline data was 10,671.1 km in length with 134 of the grids being more complete than the Tiger data. The Tiger data was 11,253.4 km in length with 163 grids being more complete than the Street Centerline data. The Street Centerline data had 5.4% less road coverage than the Tiger data.

After assessing the data in Excel, I returned to ArcGIS and created a field to calculate the percentage difference between the datasets. There was an median of -0.18% difference in coverage between the Street Centerline and Tiger data.

I used this percentage difference field to create my choropleth map:





Tuesday, August 27, 2024

Special Topics- Data Quality- Standards- Module 2

 This week we used the standards set by the National Standard for Spatial Data Accuracy (NSSDA) to determine positional accuracy of a provided road network. The positional accuracy of the test data was compared to an established data set as well as to reference points obtained to orthographic images of the study area.


To determine accuracy we chose a sampling of 20 points meeting specific criteria (no less than 4 per quadrant, greater than 10% of the diameter in distance from other points) along intersections. We established reference and testing feature classes for each intersection. 





After creating our feature classes we added XY coordinates to the data. We used excel to calculate the RMSE and NSSDA values.

Based on my results I determined the following about the positional accuracy of each data set:

The ABQ street data tested  4.18 feet horizontal accuracy at 95% confidence level.

The Street Maps USA data tested 434.85 feet horizontal accuracy at 95% confidence level.


Wednesday, August 21, 2024

Special Topics- Calculating Metrics for Spatial Data Quality- Module 1

 This module covered spatial precision, accuracy and bias. We also learned how to calculate root-mean-square error and cumulative distribution function.

Accuracy refers to the closeness of a collected data to an accepted reference point, while precision refers to the agreement between the data points. 

To determine accuracy we created a feature class with a single feature that represented the average of the latitude and longitude of all the collected data. We then compared this with a provided reference point that represented the "true" value.

To determine precision we spatially joined the average point feature class with the original point data feature class to obtain a distance attribute calculation. We calculated the distance of 50, 68 and 95 percentiles of the data from the average location, and used the buffer tool to create a visual representation of the percentiles. After determining the horizontal precision and accuracy we then determined the vertical precision and accuracy using the altitude data.

Our horizontal accuracy was 2.97 meters while our horizontal precision was 4.17 meters. 

Sunday, August 4, 2024

Module 6- Least Cost and Corridor Analysis- Applications in GIS

 The second part of our final module in Applications covered least cost path and corridor analysis using cost surfaces. We worked with two scenarios. For the first scenario we analyzed slope, river crossings and river proximity to determine the best single possible location for a new pipeline within a given study area. We then took this information a step further and presented a corridor analysis. To do this we reclassified our features with their cost factor and utilized the Cost Distance and Cost Path tools.



Image 1


First we assessed for slope alone. As you can see this first image contains 4 water crossings. 

Image 2


After adding the addition cost value of the river crossing the adjusted pathway only contained two water crossings (yellow). However, we further adding a 3 raster feature layer to the cost analysis for river proximity. This gave us the final path in image 2 (red).
Image 3

    Finally, we used the corridor tool to get a corridor analysis of the data. Image 3 depicts a copy of the final least cost analysis overlaying the corridor analysis



    After establishing the basics of performing a corridor analysis we were tasked with determining the most likely corridor for black bear migration between two protected areas of Coronado National Forest. We evaluated suitability based on slope, proximity to roads, and landcover surveys. 
    After giving suitability ratings to each of my variables, I ran a weighted overlay to establish the importance of each variable for the cost surface. I ran my cost surface in the cost distance tool with each protected area to create two rasters for the corridor output. When running the corridor analysis I got a very widespread result that I wanted to narrow down. To do this I created an empty polygon feature class for my proposed corridor and drew a polygon around the most suitable of the corridors. I then clipped my raster to my proposed corridor.

Here is my final map:


 




GIS Portfolio

 We were tasked to create a GIS portfolio for our internship program. It was a great opportunity to put organize the work I have been doing....