GitHub - qichenftw/dpark: Python clone of Spark, a MapReduce alike framework in Python

DPark is a Python clone of Spark, MapReduce(R) alike computing framework supporting regression computation.

Example for word counting (wc.py):

 import dpark
 file = dpark.textFile("/tmp/words.txt")
 words = file.flatMap(lambda x:x.split()).map(lambda x:(x,1))
 wc = words.reduceByKey(lambda x,y:x+y).collectAsMap()
 print wc

This scripts can run locally or on Mesos cluster without any modification, just with different command arguments:

$ python wc.py
$ python wc.py -m process
$ python wc.py -m host[:port]

See examples/ for more examples.

Some Chinese docs: https://github.com/jackfengji/test_pro/wiki

DPark can run with Mesos (0.9 or latest).

If $MESOS_MASTER was configured, then you can run it with mesos just typing

$ python wc.py -m mesos

for shutcut. $MESOS_MASTER can be any scheme of mesos master, such as

$ export MESOS_MASTER=zk://zk1:2181,zk2:2181,zk3:2181/mesos_master

In order to speed up shuffing, should deploy Nginx at port 5055 for accessing data in DPARK_WORK_DIR (default is /tmp/dpark), such as:

        server {
                listen 5055;
                server_name localhost;
                root /tmp/dpark/;
        }

Mailing list: dpark-users@googlegroups.com (http://groups.google.com/group/dpark-users)

Name	Name	Last commit message	Last commit date
Latest commit History 481 Commits
dpark	dpark
examples	examples
tests	tests
tools	tools
.gitignore	.gitignore
AUTHORS	AUTHORS
CONTRIBUTORS	CONTRIBUTORS
LICENSE	LICENSE
README.md	README.md
TODO	TODO
setup.py	setup.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

About

Uh oh!

Releases

Packages

Search code, repositories, users, issues, pull requests...

License

qichenftw/dpark

Folders and files

Latest commit

History

Repository files navigation

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Packages