2-parts. features a ton of interviews with high level executives including the facebook "head of growth"
and alex stamos who was the security chief. I thought stamos came out of the interview seeming pretty forthright and direct (he no longer works at facebook). the other executives did not come off so well...
I don't know if THIS article was directed by Facebook related PR but the times reporting says there is other PR basically identical to this. So I think we can decisively label your conjecture as "credible"
how would hear about this story unless you are personally friends with alex stamos or zuckerberg or sandenberg? I am all for skepticism but blanket rejection of responsible journalism seems like an over-reaction. the reporting of the new york times on facebook has been continuously borne out by events.
1.mmmmmmmmm ok I am willing to accept you meant the quadratic loss instead of 0-1 error. that seems reasonable.
2. this is paper is centered in a research thrust that IS focused on generalization. see my below comment.
I don't know who most people are but this paper COULD be important in understanding why stochastic gradient works well in practice.
Personally I doubt it very much.
3. massively overfitting to the training dataset BUT
generalizing well is a real phenomenon and yes it is very weird. happens in deep nets and i believe adaboost. i.e. continuing to train after you have zero 0-1 loss. I agree this is a weird way to communicate this idea but that is what the community uses.
have you heard of something called ERM? uniform convergence?
The typical way of showing generalization in ML is to show that if we have some low or zero error solution on the test data-set, for a large enough dataset, with high probability, the error on our training data set is close to the error on the real and unknown distribution. The first step which is basically "find a low error hypothesis on the training data" is called the ERM principle.
In practice we observe stochastic gradient descent works pretty well in solving the ERM problem and the solutions generalize well (perform well when deployed).
This is very weird since neural networks are really weird objects with very non-linear and non-convex behavior and gradient descent shouldn't play well with weird bumps and curves and valleys.
People want to show mathematically that stochastic gradient descent does well on neural networks.
This paper claims gradient descent is effective at minimizing quadratic loss on the training data.
If we could improve the results to show that on the true distribution we also have low loss-that might be compelling that gradient descent converges to the minimum error solution.
None of this explicitly stated since this is a well understood part of basic literature in learning theory.
Showing an algorithm can do erm on the hypothesis class is the first and (easier ) part of showing generalization.
If you want a good reference that explains this in a more coherent way I recommend looking at the first 4 chapters of understanding machine learning theory by Shai-Shalev Schwartz.
If you still think the comments I was responding to are not totally incoherent-take note of the fact that the very first sentence in the paper is "One of the mysteries in deep learning is random initialized first order methods like gradient descent achieve zero training loss"
1. people overfit the baby datasets to zero training loss (MNIST) all the time. maybe you meant a "hard" dataset.
2. You clearly have no idea what you are talking about.
This paper is trying to argue a bit about why neural networks generalize well by showing with math that a nn with some of their conditions converges to the zero training loss. It isn't remotely meant to be practical. IT IS A THEORETICAL PAPER.
And comparing it to nearest neighbors of 1 is so so so so so silly it isn't even wrong.
edit. #1 is actually an entire research direction in the theory of machine learning fyi.
It is possible to get neural networks that massively overfit but still generalize (which Is weird).
> Well, maybe because it's all done by a priest-caste cartel, shielded from reality in ivory towers, and then presented to the public in a similar way religion once was forced on peasants (taxes included). Academia is the modern priesthood: usurping the right to all knowledge, in bed with the state, corrupted to the bones, fighting heretics.
I don't mean to be rude....I actually fuck that. I mean to be rude. You are an irrefutable argument against public-review.
There is a serious and well documented basis in neutral american and european media-going back decades.
It reminds me of how so many smart people make decisions based on an anecdotal basis, including American media like NYT.
I implore people to not do that again.
side note: your background is wild. you have been around the block a time or two..I think you have some good stories to tell.
https://www.pbs.org/wgbh/frontline/film/facebook-dilemma/