MapReduce is a framework introduced by Google for processing larges amounts of data.
The framework uses a simple idea derived from the commonly known map and
reduce functions used in functional programming (ex: LISP). It divides the main problem into smaller sub-problems and distribute these to a cluster of computers. It then combines the answers to these sub-problems to obtain a final answer.
MapReduce facilitates the process of distributed computing making
possible that users with no knowledge on the subject create their own
distributed applications. The framework hides all the details of parallelization,
data distribution load balancing and fault tolerance and the user
basically has only to specify the Map and the Reduce functions.
In the process, the inp
ut is divided into small independent chunks. The map function receives a piece of the input,
processes it, and passes the input in the format key/value pair as
answer. These key/values are grouped in a certain way and given as input
to the reduce function. This in its turn merges the values, giving the
final answer.
Each map and each reduce may be processed by a different node (Computer)
in the cluster. The quantity of nodes in the cluster may be as big as
the availability of computers in your network. The framework is
responsible for dividing the input and feeding the map function.
Afterwards, it collects map's outputs, group and send them to the
reduce function. After the work of reduce if done, the framework gather
the answers in a final output.
The following picture shows the MapReduce flow