Well then it shows that these models are using widely disparate training sets and have high confidence even when they shouldn't.
Questions like "is mouthwash effective" presumably has one solid data source -- medical journals.
Questions like "is mouthwash effective" presumably has one solid data source -- medical journals.
You can argue the study isn't as case-closed-decisive as we'd ideally like, but it's certainly evidence. It's probably hard to design a better study.