True in the general case, but the idea is to set things up (possibly by constantly changing the environment) so that there isn't one. This has happened at least once, when the brains of our ancestors have tripled in size relatievly quickly.
There is no question people were trying related approaches in the past (as the paper cites!). Indeed, it's exceedingly rare to find a totally new idea; we always stand on the shoulders of giants, which include but is not limited to Steels's work. The goal, both here and in general, is to go beyond previous work — in the case of this work, to use modern DL tools to learn a more sophisticated language that has a real degree of compositionality and grounding. The hope is that by pushing this approach to its logical limit, we'll get agents that can really understand language, both artificial and natural.