Version 1 and Version 2 use exactly the same dataset. Version 1 is milder and easier to handle, while Version 2 is trained almost to the edge of overfitting, giving it a much stronger and more distinctive style.