But, when you revise the model based an analyzing the predictions of the previous run and adapt to the positive result, isn't that supervised learning with back-propagation?
I would agree with this approach. In evolution, it would be the equivalent to a conditional lethal mutation. In humans and even non-humans, may behaviors are learned through a form of adaptive behavior that oftentimes becomes a form of abductive reasoning. What is learned is "good enough" to serve as ground truth until there is evidence to contradict.