GeoSparkViz: A scalable geospatial data visualization framework in the apache spark ecosystem

Jia Yu, Zongsi Zhang, Mohamed Elsayed

Research output: Chapter in Book/Report/Conference proceedingConference contribution

1 Scopus citations

Abstract

Data Visualization allows users to summarize, analyze and reason about data. A map visualization tool first loads the designated geospatial data, processes the data and then applies the map visualization effect. Guaranteeing detailed and accurate geospatial map visualization (e.g., at multiple zoom levels) requires extremely high-resolution maps. Classic solutions suffer from limited computation resources and hence take a tremendous amount of time to generate maps for large-scale geospatial data. The paper presents GeoSparkViz a large-scale geospatial map visualization framework. GeoSparkViz extends a cluster computing system (Apache Spark in our case) to provide native support for general cartographic design. The proposed system seamlessly integrates with a Spark-based spatial data management system, GeoSpark. It provides the data scientist a holistic system that allows her to perform data management and visualization on spatial data and reduces the overhead of loading the intermediate spatial data generated during the data management phase to the designated map visualization tool. GeoSparkViz also proposes a map tile data partitioning method that achieves load balancing for the map visualization workloads among all nodes in the cluster. Extensive experiments show that GeoSparkViz can generate a high-resolution (i.e., Gigapixel image) Heatmap of 1.7 billion Open-StreetMaps objects and 1.3 billion NYC taxi trips in ≈4 and 5 minutes on a four-node commodity cluster, respectively.

Original languageEnglish (US)
Title of host publicationScientific and Statistical Database Management - 30th International Conference, SSDBM 2018, Proceedings
EditorsMichael Bohlen, Johann Gamper, Peer Kroger, Dimitris Sacharidis
PublisherAssociation for Computing Machinery
ISBN (Electronic)9781450365055
DOIs
StatePublished - Jul 9 2018
Event30th International Conference on Scientific and Statistical Database Management, SSDBM 2018 - Bolzano-Bozen, Italy
Duration: Jul 9 2018Jul 11 2018

Publication series

NameACM International Conference Proceeding Series

Other

Other30th International Conference on Scientific and Statistical Database Management, SSDBM 2018
CountryItaly
CityBolzano-Bozen
Period7/9/187/11/18

Keywords

  • Big spatial data
  • Distributed computation
  • Spatial visualization

ASJC Scopus subject areas

  • Software
  • Human-Computer Interaction
  • Computer Vision and Pattern Recognition
  • Computer Networks and Communications

Fingerprint Dive into the research topics of 'GeoSparkViz: A scalable geospatial data visualization framework in the apache spark ecosystem'. Together they form a unique fingerprint.

  • Cite this

    Yu, J., Zhang, Z., & Elsayed, M. (2018). GeoSparkViz: A scalable geospatial data visualization framework in the apache spark ecosystem. In M. Bohlen, J. Gamper, P. Kroger, & D. Sacharidis (Eds.), Scientific and Statistical Database Management - 30th International Conference, SSDBM 2018, Proceedings (ACM International Conference Proceeding Series). Association for Computing Machinery. https://doi.org/10.1145/3221269.3223040