There seems to be two factors at play here that are creating almost nonsensical results.
1. The results, when filtered through the replicability guidelines rendered a clear verdict - 61 to 39 against.
2. The question of "how closely did the findings resemble the original study?" flips the findings. Moderately similar findings are the majority - 58 to 42 in the other direction.
How can you have a study that has "virtually identical" findings that doesn't replicate the original?