Humans are not very good at understanding large quantities. With
increasing media reports about “Big Data”, this problem is even
more apparent. Descriptions of Big Data frequently refer to its
“Three Vs”: Volume (storage required), Velocity (speed of data
acquisition), and Variety (complexity of data types).
How big is a megabyte? Perhaps about the size of a large novel
(400-500 pages or so). Easy to grasp (literally!). How big is a
petabyte? Much harder to imagine! Reporters and some data scientists
often try to translate such enormous numbers into familiar sized
chunks. Like a recent article about melting ice in Greenland
(https://www.cnn.com/2022/07/20/world/greenland-heat-wave-ice-melting-climate/index.html)
using MSPDs (millions of swimming pools per day). Not very helpful,
really, although the article did later translate that into a “foot
of water covering West Virginia”, a bit easier to visualize given
the recent news reports of flooding in Kentucky.
When describing Big Data sizes, avoid using unusual units, like
swimming pools, ten dollar bills laid end to end from the earth to
the moon, or other clever but basically useless metrics. Keep it
simple and visual. For example, if a byte is like a single grain
of rice, a cup of it is about a kilobyte (KB), a gigabyte (GB) is
a few truckloads of the stuff; a zettabyte (ZB) would fill the
Pacific Ocean.
Think about another big data V: Velocity. Last year (2021) Twitter
averaged about 200 thousand tweets per minute. Visualize every
person in Ohio State University’s huge stadium continuously staring
at their phones and tweeting instead of watching the football game.
Try to visualize the mind boggling amounts of data you are generating
and consuming every day with your Internet and social media activity.
Then think about what system engineers and analysts must do to
extract useful knowledge from all that content. That remains today’s
and future’s rapidly growing challenge of Big Data. And some great
employment opportunities!
