Yes, but you can certainly automate the process to a large extent. For example, many sites simply don't have the functionality allow copyright infringement. Also many sites will have >99% false positives, whereas torrent sites will have >99% true positives.
They could hire a few people to get through as many computer sorted reviews as possible, and the rest of them fall through the cracks.
The only reason to use polymer is if you need the L style UIs they discussed at Google I/O.
While Google's really good at Java apis, they're really terrible at JS apis. I love Google, but I would be very hesitant to use their JavaScript libraries or frameworks. A lot of people I've talked to (on irc, es-discus, etc.) don't agree with web components, but Google's pushing it hard.
Web components are poorly based good ideas, and polymer is poorly based on web components, and your app will be (poorly) based on that.
Other people have started posting some of the problems with polymer.
StackOverflow, Wikipedia, and other creative commons sites provide good starting data sets because all you need to do is attribute to them. No expensive licensing things, or anything like that.
Often you can write a pretty good bot using their data, and emulate user interactions. Enough users won't notice the site is run by bots that it's better than nothing if you don't have funding :-)
They could hire a few people to get through as many computer sorted reviews as possible, and the rest of them fall through the cracks.