It also resonates especially when you're reading from hard experience. A year ago I would have quickly read & enjoyed this, then moved on. Now too many items give me pause as I recall some mistake I've made. It is very good.
Can you tell apart founders who sponge up advice vs those who are doomed to learn on their own? I suspect the answer is conversational resourcefulness (http://www.paulgraham.com/word.html) but wonder if you have more to say.
What he said! I'd like to add that a key to squeezing more out of NathanRice's post is the phrase "conjugate prior." Another totally natural thing would be to use a Gaussian prior & likelihood, then update the posterior as ratings arrive. This would take advantage of the ordinality of ratings as NR suggests. Bishop's Machine Learning book goes into this sort of stuff in more depth.
Hey, AOLserver is open source. At the time it was the only web server with a built-in glue language (Tcl) and pooled database connections. It got me excited about web programming. Plus, the sheer amount of code already written was fantastic.
I never understood why people were put off by his favoritism to those tools. ACS was all about page flow and user experience, and in that department, got it right more often than not.
Kevin Garnett plays exactly the same way that you describe Jack Lambert. He goes only one speed: full blast. Whenever I need motiviation, I find myself asking myself, wwkgd?
Creating the list of URLs and prioritizing them is the hardest thing about building a crawler! That is, a good, web-scale one. A replacement for wget might be sort of fun, but the real way to make a fast crawler is to be choosy about which pages get updated frequently, which are likely to contain good content (by computing a pagerank-like stat on the fly), etc.
It is far from my area of expertise, but the Wikipedia page about this looks very useful. It cites a bunch of wicked smart people. http://en.wikipedia.org/wiki/Web_crawler
If you just want to suck down a bunch of pages, then there's nothing wrong with wget.
Sure, there's plenty of fluff in the article, but it isn't because the author (http://www.math.wisc.edu/~ellenber/) is clueless about science. The article is probably at the right level of technical depth. Seems ambitious to even touch SVD in Wired.
Nespresso is great. At $0.50 per pod, it's pricey, but for such painless cleanup and consistently good shots, worth the price. Hands down the best teas I've had come from Upton Tea (uptontea.com).
My intuition says it's not possible to get 1 without destroying the ability to get 2, or vice versa. But I'm no expert. The solution that leaps to my mind is to use a locality-sensitive hash function. It would make search stochastic, which means reversing the encoding would be easier. Seems like any hash-based solution will involve a tradeoff between 1 and 2.
That's not true. You can submit your results and then refuse to release your algorithm, disqualifying you from the competition.
"Upon qualifying, as described above, the Participant is required to submit within one (1) week for judging a description of their algorithm along with all source code. The Participant warrants that the source code is either fully or substantially developed and functions or will function as represented by the description. Failure to deliver both the description and source code within one (1) week will disqualify that entry and additional qualifying entries will be considered."
The non-friend vs anti-friend distinction didn't even occur to me, but it's clearly an important part of the experiment. I like all the discussion of data. Fun problem.
I did not realize squaring an adj matrix tells you what it does. Thanks for edjumacating me.
Did you go past f2hops? Seems like 3 would be reasonable and predictive. Since 1/2 of your tree is so small, and there were 12k nodes in the tree, that suggests to me a pretty easy task. Do you agree? It would be interesting to see if PCA or LDA pick the same features as the decision trees did. Just a click away in Weka, after all.
(An aside, and neat hack: Buddy of mine just walked in and saw the document on my screen. He saw the decision tree and said, "I remember those. In grad school I printed out decision trees as C if/else statements. Part of running my decision tree was a call to gcc.")
Yeah, yeah, so the engineering part wasn't that great. But were the results good? Which algorithms did you use? What were your features? Seems like an interesting experiment, especially with all the gratuitous "friending" people do. Reminds me of a related paper: http://www.hpl.hp.com/research/idl/papers/facebook/facebook.pdf
Good cites. I second them. Mackay and Jaynes are excellent but can get out of hand quickly. Duda, Hart and Stork is another good hardcore option. When I need the book "for dummies," which is pretty often, the Weka book and Mitchell's Machine Learning are helpful.
When you file yourself (pro se), the PTO is gentler and kinder than they are if a lawyer files for you. If you have the time and energy, file now to get the early filing date and call a lawyer later.
Can you tell apart founders who sponge up advice vs those who are doomed to learn on their own? I suspect the answer is conversational resourcefulness (http://www.paulgraham.com/word.html) but wonder if you have more to say.