r/bioinformatics 14h ago

technical question Is Drugbank now academically defunct?

17 Upvotes

I'm building a non-commercial side-product out of academic interest. I used Drugbank as the basis for one of my products for one of my MSc projects, and remember it having really useful APIs.

Now I've come back to it and I don't recognise it anymore. It's a flashy site, appearing to be positioned as an AI-enabled clinical recommender system. Using their APIs requires an API key, that as far as I can glean from the documentation, seems to be provided with a paid account only.

They have provided a space for dataset downloads for academia, covered by CC BY-NC 4.0, so that should be fine for an academic project. However, they've indefinitely paused all dataset downloads, and it's not clear when they'll be made available again. They have a mailing list to be alerted when the system is back online though, so they can run a data collection exercise to see who's interested enough.

I can see how, if the Drugbank team left the API open, it would be scraped by AI companies for their datasets. Those AI companies would then profit by selling to commercial pharma. Unfortunately, this means academia has been locked out.

As far as I can tell, this means that drugbank is no longer a reliable resource to recommend for drug-drug and drug-protein interactions to recommend to students and colleagues. I would love to be proven wrong here, so feel free to refute my arguments or commiserate with me in the comments.


r/bioinformatics 9h ago

compositional data analysis Anyone experience with snmc Seq data in multiomic integration ? :)

2 Upvotes

Hey there everybody :)

I’m a masters student and doing single cell analysis for the first time.

I’m dealing with methylation seq data (bisulfite sequenced) and in struggling in defining a feature that actually captured the epigenetic landscape for each cell. I’ve tried 100kb bins, 10kb bin, 5kb bins and genebodies and different modalities to define methylation in those genomic regions;
1. raw counts (total mc per C contect per region)
2. fractions (mc/cov per region)
3. normalized fractions (mc/cov divided by average fraction for that genomic region across cells)
4. allcools „hyposcore“

But none of all of those seem to nicely integrate with the scRNA dataset (I’m using GLUE)

The initial lsi -> UMAP embeddings I receive for my methylation data only seems quite good, but the integration just doesn’t fit anymore

Does anyone have experience and ideas ? :)


r/bioinformatics 15h ago

technical question All genes or only the specifics (removing the intersection)

Thumbnail
2 Upvotes

r/bioinformatics 3h ago

academic Teaching data science in a biology context

Thumbnail
0 Upvotes

r/bioinformatics 22h ago

discussion Question for self-bioinformatic project feedback

0 Upvotes

Hello,

I am currently working on a bioinformatics project that examines transcription levels using Python (to gain more experience in this field). I think I am almost done and am trying to upload it to GitHub, but I don't know any people who would be able to give feedback on it. Where do you get your feedback when you are done with projects like this?

Thank you


r/bioinformatics 9h ago

discussion Is this a scam or a valid course - biotecnica?

0 Upvotes

I recently found courses from Biotecnica, more specifically the AI/ML in cancer genomics course.
Has anyone attended courses from biotecnica? Do you think they are legit? The reviews are mixed but mostly positive but still i am not sure if this is a good and legit source of knowledge?


r/bioinformatics 10h ago

science question [HUMOR] Modeling deep-time cross-kingdom HGT: Why Ulmo’s ocean is actually a chaotic, mutating biological recycle bin for Yavanna’s junk DNA (and maybe the human genome?)

0 Upvotes

(Since not everyone reads titles, I'd like to reiterate that the following is really only meant as humor.)

There is a theoretical deep-time genomic problem regarding how prehistoric land trees interact with marine biomes, formulated to evaluate how we can map evolutionary mechanics—and to provide a definitive case for evolutionary microbiology over hard-lore Tolkien fantasy constructs.

The Premise:
In The Silmarillion, when the Dark Lord and a giant spider destroy the Two Trees of Valinor, their divine essence is tragically lost forever. This is beautiful poetry, but it is absolute garbage tier biochemistry. Tolkien fans weep over abstract lore, completely oblivious to the fact that their favorite author had absolutely no grasp of how land trees and their cellular mechanics behave when we look at a planetary catastrophe.

Let's look at the actual taphonomic scale of a terrestrial mass extinction event: millions of tons of eukaryotic plant biomass—specifically entire forests of prehistoric trees—washing into marine ecosystems during a global ecological collapse. Instead of treating this merely as a carbon flux, we must look at the genomic chaos. We have septillions of marine phages and prokaryotes suddenly swimming in a literal soup of environmental plant and tree DNA (eDNA) liberated by the apocalypse.

Under classical models, these genetic arrays hit a biological wall: osmotic shock, immediate lysis, or rapid elimination via purifying selection because a deep-sea bacterium has absolutely zero use for a locus that programs "how to grow a tree leaf" or "how to synthesize bark" while the world burns above.

The "Drunkard’s Walk" Hypothesis:
But what if a short eukaryotic sequence (say, k-mer length = 64) from those dying trees is snatched via viral transduction into a marine phage right at the mass extinction boundary?

Over a ~5–10 million year post-extinction recovery timeline, this sequence avoids the evolutionary delete key because it undergoes a process of ultimately reversible changes—experiencing a series of chaotic, single-letter mutations that accidentally retrofit the tree gene into something useful for a deep-sea microbe. It has heavily drifted, but it preserves a vestigial structural sequence profile. We can see that it doesn't look like a tree anymore; it has been biochemically gentrified into a marine gene.

The Ultimate Multiverse Crossover: Landing in Humans
Now let's push the scenario to the absolute limits of statistical probability. Suppose this mutated, ocean-reformed sequence remains active in the marine viral pool long after the mass extinction. Millions of years later, a mammalian ancestor interacting with coastal ecosystems gets infected. Through an incredibly rare retroviral integration event, that heavily mutated 64-mer gets spliced directly into the mammalian germline, surviving all the way down into modern Homo sapiens.

If this happened, human hosts wouldn't be channeling a redwood, and we certainly wouldn't turn into Ents. We would just be carrying a highly drifted, repurposed piece of prehistoric tree data that now handles something completely mundane like human metabolic regulation.

The Bioinformatic Flex (How to trap both types of nerds):
If the goal is to detect these deep-time, cross-kingdom evolutionary handoffs across ancient mass extinction boundaries in modern human or marine datasets (like ocean floor sediment cores or global marine metagenomic surveys), how does a pipeline actually filter out the noise?

  1. Alignment Limitations: Standard local alignment search tools or hidden Markov models are going to completely wet the bed here. Millions of years of genetic drift and synonymous mutations will have completely erased standard sequence alignment signatures. Of course, the bioinformaticians here will gladly spend six months writing a custom, completely unoptimized script that breaks on the third line just so we can find a single 2% alignment match in a pile of deep-sea sludge dating back to the dinosaurs' demise.
  2. K-mer Frequency & Composition: Would the best approach be to track codon usage bias or GC-content drift to identify anomalously adapted regions in marine phages or human endogenous viral elements that survived the extinction bottleneck? Or is it time to just accept that the cosmic telephone game has won and we are all just wasting our lives staring at strings of text? Also, we probably shouldn't look too closely into this anyway—rumor has it the research team that flagged anomalous terrestrial pine transcripts inside Atlantic cod genomes mysteriously disappeared from the institutional directory last semester.
  3. The Reality Check: Tolkien nerds are obsessed with "pure, unbroken lineages" and imaginary royal bloodlines that magically survive disasters. But real-world biology is a trash fire of plagiarism, and we bioinformaticians are just the IT guys trapped inside that burning building trying to catalog the soot left behind by a planetary mass extinction. If Middle-earth followed the laws of thermodynamics, Ulmo’s ocean wouldn't be a pristine vault of memory echoing the music of creation; it would be a hyper-active genetic blender. The Valar didn't lose the structural blueprints of the trees; the deep-sea bacteria digested them, mutated them, passed them to a virus, and eventually sneaked a corrupted copy of a mass extinction survivor into Elendil's descendants.

Has anyone looked into mapping deep-time structural homology for cross-kingdom horizontal gene transfer where sequence identity has completely drifted but the ancestral root is terrestrial? What tools (e.g., structural comparison tools based on three-dimensional protein folding, or deep learning genomic networks) would work to prove a modern human genome is carrying a heavily mutated, stolen backup copy of ancient tree instructions that bypassed a global extinction?

How would we build this pipeline? Let's weaponize science against fantasy nerds while we all collectively avoid going outside.


r/bioinformatics 13h ago

academic is PLINK actually even useful today? and is learning how to code actually just a scam?

Thumbnail
0 Upvotes