Edit: to elaborate, certain amino acid sequences are seen again and again, and others are practically impossible to see in nature because amino-acids fold into 3D proteins...and each amino acid is like a Lego piece. There is a lot more detail needed to understand fully how tertiary and 4try protein structures form...but it is also possible to understand it decently well after an undergraduate degree in biology, chemistry, biochemistry, etc. A PhD in the field knows this inside and out. If I had to guess, the author of this post has a PhD in something like statistics or computer science, and thinks that they can apply high school level math to two fields (molecular biology and biochemistry) that they do not even have a high school level of understanding of.
To make a very stretched analogy (I am a doctoral student in the life sciences and a hobby programmer), this blog post is like saying that some Java source code is stolen because four of the class names are the same between two projects, and then using the number of possible characters in the class names and the length of the name to run some statistics. The problem being that no one has class names like "AeNOQ92bA"...in fact the vast majority of 8 character sequences will never be class names. Just like the vast majority of amino acid sequences (likely) do not exist in nature. And then you dig deeper and find out the class name is something like "MainDashboardSupervisorTree" (I do not know Java so forgive me). Then you also find out that the author of both classes...is the same guy who moved companies and likes particular naming convection, but never meant to copy stuff word for word. Similar to how SARS-CoV-2 could have naturally incorporated HIV-1 RNA into its genome when co-infecting a host.
> It’s easy enough to change a single nucleotide (a single point mutation or SNP) or even insert or delete nucleotides (less common) but to insert 20 or 30 nucleotides with a code that works? Nope, that has to come from another virus or else it’s been done in a lab.
The original preprint which noted similarity of the HIV-1 sequences and those in the COVID spike protein, and on which this blog author's claims rest, was voluntarily withdrawn by its authors. [1]
Moreover, the assertions of uniqueness have been thoroughly addressed in the published article, "HIV-1 did not contribute to the 2019-nCoV genome" [2] :
> For any virus to obtain additional insert sequences from other organisms, it requires that it has direct interactions with other organisms, most likely through homologous or non-homologous recombination. [...] On the contrary, these motifs are widely present in various mammalian cells and so it will be more likely for bat CoV viruses to gain those motifs from the genomes of their infected cells if recombination indeed occurs.
That rebuttal article clearly suggests several feasible sources of the gene in Nature. That surely is more plausible than suggesting that Moderna or some other group nefariously created the COVID genome _in vitro_ before unleashing it upon humanity.
I deeply enjoy Hacker News for its ability to surface different ideas and sources of information that I would otherwise never find, but this article is simply ridiculous.
[0] https://en.wikipedia.org/wiki/Horizontal_gene_transfer
[1] https://www.biorxiv.org/content/10.1101/2020.01.30.927871v2
[2] https://www.ncbi.nlm.nih.gov/labs/pmc/articles/PMC7033698/
This would be a particularly useful discovery, because it would dramatically limit the candidates for who did the gain of function research to organizations licensed this Moderna lineage.
There also seems to be some circular reasoning the argument. Apparently we can ignore RaTG13 because it’s obviously synthetic, which makes SARS-CoV-2 look even more synthetic. It would be interesting to compare to the BANAL family of SARS-CoV-2-related viruses that are even more closely related to SARS-CoV-2 than RaTG13 [1].
I’m not sure why only viral genomes were searched for the furin cleavage site sequence. Viruses famously exchange genetic material with their host organisms. The “smoking gun” sequence also appears Mycobacterium smegmatis, for example [2].
[1]: https://www.nature.com/articles/d41586-021-02596-2 [2]: https://twitter.com/soychicka/status/1243547603746410500
Let's try a more sensible approach. Let's see how often the sequence occurs in other databases. The author's 1/20^6 (1/64,000,000) probability calculation omits the fact that every database search considers many millions of possible alignments (a crude calculation would be 6 * the database length, or 6 * 1.6E7/20^6 ~ 10000E6/64E6 = 156 times by chance), and we're shown that in a 5,000,000 protein/1.6E7 residue database, 100% matches are seen several dozens of times, as expected.
How about other protein sets? Search the human protein set (19,000 proteins), there are no 6 amino acid 100% identical matches. The best matches are 100% of 5 residues, or 83% identity over 6 (with an E()-value of 22).
How many proteins do we need to search to find an exact full-length match? (we have dozens of matches in a 5,000,000 viral protein database. and none in 19,000). If you search all the NCBI landmark sequences (~500,000, non-redundant), there are 3 identical matches (E()-value: ~44) in Drosophila, Corn, and Strep. pneumoniae. So this sequence is not very special -- as the E()-value statistics indicate, it happens all the time by chance. (If it happens 3 times in a 500,000 sequence non-redundant database by chance, we would expect it to happen 30-100's of times in a 5,000,000 redundant viral database by chance, which is what we see.)
Six amino acid non-significant matches are evidence that properly done statistics are accurate, and pretty much nothing else.
BLAST does not contain all genetic material from all viruses or organisms in all of the natural world. There are countless (really, impossible to count) viruses of all kinds out there dancing around in the bodies of all kinds of creatures. There are literally hundreds of trillions of viruses in your body right now (not different kinds but individual viruses). Each of those hundreds of trillions of replications was an opportunity to mutate. Now multiply that by all of the individual animals human or otherwise who could harbor coronaviruses and you have many hundred billions of trillions of chances to generate the gene sequences in question. If I could get billions of trillions of lottery tickets, I'd win the lottery, and here it looks like SARS-CoV-2 won the "does a gene sequence in the virus exist in a patent by moderna" lottery.
That kind of worried me. Because even if you think sars-cov-2 is completely natural (and I have no problem accepting that position), have you thought about how many groups around would still have the capacity to make such a virus, if they wanted to?
Even just the fact that "some people believe the virus was made in a lab" can be dangerous. Because what if, say, Kazakhstan's dictator-emeritus Nursultan Nazarbayev believed it? How far-fetched is it that he, or someone like him even in the more nominally democratic parts of the world, would decide, "they fired the first shot, we can't allow there to be a bioweapon gap!" and commandeer the institute to do "dual use" research?
It would be easy to excuse too. They could say, "we're just designing these powerful virus variants to have a vaccine ready in case someone else comes up with the same powerful virus variant", and it would be a reasonable argument, at least as arms race escalation arguments go.
Dismissing everything as far I'd problematic. If you can't prove it and several scientists were baffled by the number of changes and we still donate an answer, you can't criticize someone from believing A while you believe B.
This is *published research*. Specifically observations 3 and 4 https://jvi.asm.org/content/jvi/82/4/1899.full.pdf
And so, fast vaccine production must be a priority. The development of drugs to treat viral infections and their manufacture must also be a priority. The ability to detect carriers and quickly shut down travel must also be a priority. And how do we enforce quarantine in populations that are resistant to it? All of these concerns are, I am certain, at the forefront in the mind of any Western government to be certain, but likely any government anywhere.
Because regardless of IF Covid-19 was created in a lab, it COULD HAVE BEEN created in a lab. And we all got caught with our pants down.
I ain't a geneticist, so it goes without saying that I don't really know what I'm doing and I'm probably using BLAST wrong, but it seems to me like (EDIT: something very close to) CTCCTCGGCGGGCACGTAG is pretty dang common even within the human genome (let alone other species), and it doesn't seem far-fetched to me that SARS-CoV-2 might've yanked (EDIT: something very close to) that sequence from one of its hosts at some point.
EDIT: I forgot to check the completeness of the matches, and it does look like the hits tend to have one or two nucleotides that don't match. Still, it's close enough that the article doesn't seem like it presents much of a smoking gun.
¹: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Get&RID=Y3TZW9W...
²: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Get&RID=Y3UU86P...
Well this is a rough start. “This information is being suppressed, which you can tell by the fact that these linked articles are still available.” The utter lack of critical thinking doesn’t make me expect much from the rest of this.
I read this before all the lab leak theories came out, and I'm glad I did because it makes them all seem ridiculous.
https://www.scientificamerican.com/article/how-chinas-bat-wo...
The rest of the interview is how to move form using a hash table for counting k-mer frequencies to a vector using a minimal perfect hash, and compute k-mer frequencies quickly using a rolling hash. It's a great question because it always throws the CS people off to hear their questions asked with biological structures involved.
Well that's just entirely false.
https://www.sciencedirect.com/science/article/pii/S187350612...
I'd draw your attention to section 2.3:
> 2.3. Furin cleavage sites are common in Betacoronavirus
Oh, also, if you click on Prashant Pradhan's paper it has been withdrawn:
https://www.biorxiv.org/content/10.1101/2020.01.30.927871v2
> Abstract
> This paper has been withdrawn by its authors. They intend to revise it in response to comments received from the research community on their technical approach and their interpretation of the results. If you have any questions, please contact the corresponding author.
1. Why do we care so much about a furin cleavage site? The author makes a big deal out of this and I'm not sure why. It seems like he specifically looked at the gene sequence at a furin cleavage site and then said: hey look, this occurs at a furin cleavage site! It must mean something! (Thanks to etaoins for finding at least some evidence that furin cleavage sites are not, in fact, unusual in Coronaviruses, although I have not read the paper: https://www.sciencedirect.com/science/article/pii/S187350612...)
2. What does HIV-1 have to do with anything? That is, if these sequences all matched virus XYZ instead, would that also be some sort of smoking gun? The top 3 sequences all seem to be very common among lots of viruses (rotavirus, measles, coronaviruses, and HIV all show up in the top 3 lists). So he seems to have found a single sequence that lots and lots of viruses have (at least variants of), and that one sequence is the first 3 rows of his table. Am I missing something?
3. If we ignore every result entered after Feb 2020, aren't we throwing out a lot of sequences that were entered specifically because people were suddenly all very interested in Coronaviruses? If he's making claims like "these don't appear in any other Coronavirus", the "except those we found after everyone started looking at Coronaviruses much more" somewhat detracts from that, doesn't it?
It feels a little bit like finding a passage "He froze, trembling with fear" and saying: the English language has 1 million words, so the probability of finding these 5 words in order is (1 million)^5. Therefore if we find it in two novels, one must have plagiarized the other. But not really though, because "he" and "with" are extremely common words, and "froze", "trembling", and "fear" are very highly correlated to each other. I think it's pretty clear his first 3 examples are all highly correlated.
His last example (and the title of this post) seem a little harder to explain to me as a layperson, but I'd note that for a 30-nucleotide sequence, there are 3^30 other sequences that are only a single mutation away. We again see Rotavirus on his list. How far away is that Rotavirus sequence, and do we think it prohibitively unlikely this sequence came from Rotavirus?
Nature favors sequences that work. Try a billion variants, keep the one that actually helps. We don't see the 999,999,999 failures.
That covid was made in a lab? Or did covid and moderna independently reach the same "conclusion" through parallel evolutionary pressures (in nature and in research)?
You have two strings, A and B, and a list of strings L. Find the longest substring that appears in both A and B but not in any element of L.
Mutation? The same way we get every new pathogen?
(eg https://i1.wp.com/www.ghgossip.com/wp-content/uploads/2020/0...)
>Posted by 'version_five'
"This sounds like one of those "bible code" things. As someone who knows nothing about DNA sequences, what is the probability that two codes could be matched? (In a birthday paradox sense where you can go looking anywhere)"
>Posted by 'whoomp12342'
"isn't it well known that one of the main reasons why we got a vaccine so quickly was because we had done a ton of research when SARS originally broke out?"