The only reason for projection here was the newness of the Volta HW (announced 2 weeks ago). The linear scaling is proven on Pascal and Maxwell HW and due to the communication library and algorithms like Block Momentum. Note: I am a MSFT employee.
Glad to see this work was updated so that the comparisons are now equivalent. I would like to see their next paper on multi-GPU. CNTK would likely do quite there as well. Note: I work at MSFT.
From the linked arxiv paper, http://arxiv.org/abs/1609.03528 this is a very interesting use of CNTK to adapt image CNN techniques to speech recognition. Surprising that CNNs worked so well on speech audio. Full disclosure: I am a MSFT employee.