Spark is not too tricky to dive into, even though you can't really take advantage unless you have a big cluster to use :)
if you want to practice data-manipulation, and a lot of the map reduce type stuff you can do with spark, I find Pandas useful for small datasets (And a lot of overlap in functionality as far as Dataframes are concerned)
For pipeline stuff, definitely take a look at Luigi, but again without a cluster it'll be less fun. Still, if you can try automating tasks with a mini luigi scheduler on your localhost, it would be good practice
This is a really interesting initiative! Thanks for sharing! Was there anything surprising you all found? Have you had a lot of feedback from underrepresented groups?
they are. back in the day they were charged with terrorism, principally. Among sabotage and other crimes. One of the rare examples of terrorism being used outside of either a wartime threat or leftist organization/labor union.
I think that the distribution of VARD is allowed (if non-commerical), but editing the source code isn't. I think this because I cannot find the source for VARD.
That's why I think it's strange it is licensed CC.
One of the problems searching in these books is that there is no standardized spelling in early modern literature.
There is a project called DREaM, at McGill to standardize for "distance reading" (macro analysis).[1] It uses a program called VARD (a text preprocessor trained to correct spelling).[2]
Strangely, this application is licensed with the creative commons. I think this means that it is closed source. Does anyone know of any open source alternatives?
It cannot handle such an immense amount of data,[3]