Well, I'm distributing trained models with this. Users shouldn't need to retrain unless they're doing research, in which case they should have access to the data.
I haven't read LDC's license on Penn Treebank recently, but AFAIR you cannot just redistribute models that were trained on the Penn Treebank. Or put differently, you can distribute the model, but any users still have to obtain a license for the treebank. That's why we are still stuck with the Brown corpus and such.
I don't understand why Google gave the English Web Treebank to the LDC. Why not just distribute it themselves?
I haven't read LDC's license on Penn Treebank recently, but AFAIR you cannot just redistribute models that were trained on the Penn Treebank. Or put differently, you can distribute the model, but any users still have to obtain a license for the treebank. That's why we are still stuck with the Brown corpus and such.
I don't understand why Google gave the English Web Treebank to the LDC. Why not just distribute it themselves?
Indeed.