I don't think this approach would scale, due to the time investment required. It also suffers from making it hard to compare one candidate to another in a fair way, unless you have everyone fix the same bugs. Having a bunch of canned bugs to be fixed doesn't seem much better than asking a CS puzzle.