Indeed. But I chose the problem in response to your comment:
> But we give no indication that the questions are hard because of computational tasks and we give many indications that the problems are of theorecical nature and hard for theoretical reasons.
and Question 034 seemed to be one of the few that did not have a computational component, and so would presumably be hard for theoretical reasons. I already indicated above some problems that I feel to be of a non-theoretical nature.
By [1, Theorem 4.1], the Neron-Severi rank of the perfectoid cover is the same as the Neron-Severi rank of the reduction. For a product E x E' of elliptic curves, it is well known that NS(E x E') = NS(E) + NS(E') + Hom(E,E'); see [2, Prop. 2.3]. Since E = E' here and E is supersingular, this number is 1 + 1 + 4 = 6.
Is it research level? It of course takes a graduate student a long time to understand, say, what a perfectoid space is. But the statement follows immediately from quoting the literature, as long as one knows what to quote.
> 1. Why do you compare it to multiplying two 1000 digit numbers and not to factorizing a 4096-bit numbers into its 2 prime factors, when not knowing any details?
The objection is to phrasing "much harder". One should distinguish between something that is difficult for reasons stemming from a lack of computational power and something that is difficult for reasons stemming from a lack of relevant abstractions or the ability to grapple with them. If the reason that a particular problem is "hard" for a PhD student is that they have to do a long calculation, but not because of a lack of conceptual understanding, then it doesn't say much about the capabilities of generative AI if the computer solves it.
Hence the example: multiplying two large numbers is hard for the former reason, not the latter. Your example of factoring a 4096-bit semiprime is hard for both reasons (because the brute force method is too slow).
I don't like that you've called these problems "research-level", or your description that they are something you might give to a second-year PhD student. Some examples:
- Question 093 is a word problem of the kind that I would imagine is commonly given to high school students. Maybe it is slightly more difficult, but it doesn't appear to have any mathematical relevance and nobody would ever give it to a second-year PhD student.
- Question 096 is something I would expect a computer to do easily by brute force, and has essentially no mathematical content other than doing a calculation. (Under what circumstance does one care about taking base 10 digits and interpreting them in base 11?). Again, nobody would ever assign this to a math PhD student, and I expect that any undergrad who knows how to code can give you this answer.
- Question 016 is the kind of combinatorial problem that one could expect to brute force with a computer (and some decently-written code) even before AI. Again nobody would give it to a 2nd year PhD student because it is too random and of no academic interest.
- There are questions like 026 and 014, about computing Hilbert series. Computing Hilbert series is a standard computer algebra task that nobody would want to do by hand before generative AI, and certainly not now.
Similar comments apply to many others. There are plenty of random-looking computational questions of exactly the type that one expects not only that computers cans solve, but should be used to solve, because nobody would ever do it by hand. None of them are research-level --- certainly not anything that would be considered publishable (before generative AI or after) --- despite the subtitle of the paper saying "research-level". And if you give them to a 2nd year PhD student I would imagine you would just be wasting their time.
I also don't like your phrasing "much harder than any exam question in any exam". If I ask you to multiply two 1000 digit numbers, the question is "much harder" than any question that will ever appear on any exam. Everyone understands the computer will do it instantly, and it doesn't demonstrate anything relevant. There is a clear regime in which one expects AI-type methods to perform better (combinatorial, calculation-based questions which can be answered using standard methods), and other regimes where one expects worse performance (e.g., proofs of statements that use abstract concepts). Why is there nothing here of the second type?
It was called Advanced Chess [1], now it's sometimes referred to as Centaur Chess. You can find online servers to play this kind of thing but it's mostly just a sort of curiosity. Computer-human pairs are well known to be better than just computers at chess though.
As you can tell from his commentary, he thinks the machine is rather stupid (and he is winning after the opening) but the machine is a much better calculator than he is and when the situation becomes more concrete he has to force a draw.
Of course, the computer on his phone is considerably worse than the best chess engines, but top chess players generally consider computers to be excellent calculators but dumb in terms of general strategy.
It's actually not that common. Players will use chess engines to analyze positions, but they will rarely play against them because it doesn't provide realistic practice for an actual chess match.
For example, it may be reasonable to play an aggressive variation against an opponent because you think they might have difficulty finding a response in time pressure, but a computer can make precise calculations in any situation, and so such a strategy almost always backfires.
What's more, chess is often abstracted at higher levels in terms of things like long term plans ("I want to put pressure on the c7 pawn and control c6") instead of concrete material gains, which allows players to make progress even in positions where there are no direct threats and no sensible exchanges of pieces. Computers when faced with such situations will know that the position is objectively drawn and so will just shuffle their pieces around aimlessly, and playing against this kind of thing is not very good practice for actual human opponents who will try to find ways to beat you anyway.
Summary of the paper for those who don't want to read it:
So basically there are two categories of "learning" involved in this sort of research, supervised and unsupervised. In supervised learning, someone gives the computer a long list of concepts and their attributes ("frog", "green frog", "jumping frog") and a set of pictures to go with each item, and feeds them into a visual-recognition algorithm. In unsupervised learning, the computer is given a concept like "frog" but then has to discover all the variations itself and get its own visual data to match.
The claim in this paper is that they have made the unsupervised learning as strong as the supervised learning. That is, they give the computer a concept ("frog"), it goes and searches through Google Books for common variations ("green frog", "jumping frog") and then uses Google image search to fetch images for each of those queries. They can then remove the obvious false positives (they test to see which images seem to screw up their learning algorithm and leave those out), and the result they get is on par with the supervised learning methods.
----------------------
In my opinion, this is only mildly interesting because Google Image Search functions based on human input anyway -- Google knows the difference between a "frog" and a "jumping frog" or even a "camel" simply because people on the internet caption such images and Google can make associations between images and their captions. Essentially, what the researchers have managed to do is outsource the work of some grad student to millions of people around the world through Google.
Of course, it could be argued that there is some sort of parallel with what humans actually do (we know what things are called because we hear other people call them that), but even if I didn't know the name of an animal I could still tell you when the same animal is in different pictures, and I can also tell you when it's jumping and what colour it is. I don't need to have someone caption the image for me to understand the broad range of situations to which the caption "jump" applies.
Although I used the word "build", the entire post is talking about the investment. Regardless of whether you consider the people who "built" it to be government employees, the funding was military funding. Things like ARPANET were Department of Defense projects.
The entire point is that the government can invest in things that it deems valuable but not necessarily likely to make an immediate profit.
The internet is an excellent example. Nobody would have predicted it would have taken off as it did, and so anyone who wanted to make an "internet" would have difficulty accumulating the capital for it. But since the military was able to build it without worrying about future profits, a large-scale and risky project was able to become a success.
The same is true of any long-term research, or any type of project that is primarily a public good. Most scientific research is not done with the explicit intention of "we want to make X", because what "X" is can't be known until you understand what is possible! But since nobody knows what is possible beforehand, investors aren't willing to invest in most research, because the simple advancement of scientific knowledge is something of public value and not private value.
I gave an example of an automated checkout, but my point was not that the jobs I mentioned could all be automated -- my point was that they are not necessary. When I say "necessary", I don't mean that they aren't necessary to increase corporate profits (salesman, which I mentioned, certainly are), I mean they are not necessary to ensure a stable and functioning society. The example I gave with automated checkouts is not to demonstrate that automated checkouts are _better_ than human workers -- human workers are and will likely be far more capable than machines at running checkouts for a long time. My point is that they _could_ be replaced, and that if society decided that increasing human liberty was a more important goal than certain small inconveniences at the checkout, they _would_ be replaced.
So yes, these people are "contributing to the maintenance of society", but there are conceivable alternatives that would give these individuals (and society as a whole) more liberty, and would only require some small inconveniences.
The anecdote you give of your startup which needs advertising to get off the ground is beside the point. Perhaps your startup is revolutionary, and perhaps it needs some advertising to get off the ground, but this doesn't change the fact that most advertising is simply misinformation. Television commercials which make use of lush landscapes and half-naked women to sell cars don't create rational consumers, and without rational consumers you cannot have a functioning market. If advertising was simply a way in which companies communicated well-reasoned facts about products, and came with a balanced analysis of a product and its competitors, then we could argue that advertising was working towards creating a functioning market system. Until then, advertising will simply favour those with the largest advertising budgets and those who are best at disinformation, and so it's difficult to argue it is contributing to society.
> Indication that considering what is and isn't necessary in society could "liberate" half the population? Yeah: every society that tried it, like the Soviet Union (hint: they'd kill people for trying to leave).
This is really cheap rhetoric -- and not even accurate. If I'm arguing for a society where people are liberated from work, why are you using the Soviet Union, a dictatorship where everyone worked all the time, as some kind of counterpoint? Not everything that contradicts the status-quo is totalitarian communism, you know.
Besides, the points I'm trying to make aren't even original. There are serious proposals that have been made for why society should move to a 20-hour work week to address things ranging from rising levels of depression to climate change. [1]
> The socioeconomic world is as good as it is because most people are employed, doing something which contributes to maintenance and advancement of society & technology.
Can you justify this? If we use the US as an example and look at the most common jobs, we find that over 4 million people are employed as salespersons, another 3 million as cashiers, followed by 3 million people employed serving food (including fast food), etc. [1]
As the population has been moved away from agriculture, and recently manufacturing, they have become increasingly employed in the service sector. These people are not creating and advancing the technology of tomorrow, they are preforming basic tasks that could easily be done by the customers themselves (eg. automated checkouts).
What's more, there is now millions of people employed in areas such as advertising and sales, where people are essentially tasked with manufacturing wants and increasing the amount of money people spend, which not only defeats the entire basis of a functioning economic system (rational consumers making rational choices), but ensures that additional numbers of people are employed in the manufacture of superfluous goods that people are deceived into buying.
Is there any indication that if we seriously consider what is and isn't necessary in our society, the working population couldn't easily be cut by half without any kind of imminent collapse?
It seems likely that someone involved with the contest is trying to use it for free advertising. Look at the "About" page on that flash app, the person apparently lives in Waterloo, and his email has "uwaterloo" in it.
There is supposed to be a new version of this contest coming up soon, although it seems to be temporarily stalled as the main developers are busy with other things.
I don't think anyone posting about how easy it is to circumvent the paywall understands the point: it's not targeted at you. The NYT knows that if they create a completely restrictive paywall they will die, so they're trying to let just enough people in and annoy them just enough to get people to pay for their service. Whether or not somebody with basic technical knowledge can bypass it is not the point. After all, they don't even ask you to register an account and yet they expect to stop you from reading x numbers of articles per month -- why would anyone expect this to be secure?
I still don't understand why Google hasn't done the exact same test with a control variable. Run the exact same test again on a domain that isn't google.com.
The Bing toolbar only sends information back to Microsoft, it doesn't decide what to do with it. For this kind of a task it would be simpler to use a packet sniffer anyway.
I think one important point you're missing is that a chess engine looking to mathematically prove a victory is very different from the ones that play against grand masters. Deep Blue may have been able to calculate 6 to 8 moves in advance, but that was after it pruned all the moves that were obviously incorrect. When you're trying to prove something mathematically however, you have to assume that even something as silly as sacrificing a queen for a pawn with no obvious positional gains is a valid move and calculate all possible branches taken from that move until checkmate some 30 moves down the road.
For example, checkers is a vastly simpler game than chess. Yet, it took 18 years of constant computation to solve it [1]. Even a game as simple as tic-tac-toe has 255,168 possible games [2]. The estimated number of chess games is 10^10^50 [3].
> But we give no indication that the questions are hard because of computational tasks and we give many indications that the problems are of theorecical nature and hard for theoretical reasons.
and Question 034 seemed to be one of the few that did not have a computational component, and so would presumably be hard for theoretical reasons. I already indicated above some problems that I feel to be of a non-theoretical nature.