Back to blog
How to Identify Introns in a DNA Sequence
Tutorial

How to Identify Introns in a DNA Sequence

By Abdullah Shahid · · 9 min read

The most reliable way to identify introns is to compare a spliced mRNA or cDNA with the genomic DNA. Aligned transcript blocks mark exons, and genomic gaps between those blocks mark introns for that transcript.

If a trusted gene annotation exists, use its transcript and exon coordinates first. Scanning genomic DNA for GT...AG alone produces many false positives and cannot reliably define a gene model.

A clean scientific illustration of mature messenger RNA blocks aligned to genomic DNA exons with looped intervening introns
Figure 1: Spliced transcript alignment identifies exons as aligned blocks and introns as genomic intervals skipped by the mature RNA.

How do you identify introns in a sequence?

Identify introns from transcript structure: select a transcript, determine its ordered exons, then define each intron as the genomic interval between two adjacent exons.

Use one of three workflows:

  1. Read exon coordinates from a trusted reference annotation.
  2. Align mRNA or cDNA to genomic DNA with a splice-aware aligner.
  3. Predict a gene from genomic DNA only when annotation and transcript evidence are unavailable.

The first two methods are much more reliable than motif scanning because they use transcript evidence and genomic context.

What is an intron?

An intron is a transcribed region removed from a precursor RNA during splicing, while the retained segments are joined to form a mature RNA.

For a protein-coding transcript, exons may contain coding sequence, untranslated sequence, or both. “Exon” does not mean “protein coding.”

Introns belong to transcripts, not simply to genes in the abstract. Alternative transcripts can use different splice sites and exon combinations, so one genomic region may be intronic in one transcript and exonic in another.

This transcript dependence is why a coordinate without a transcript accession or stable identifier can be ambiguous.

Can you find introns by looking for GT and AG?

You can use splice-site motifs as supporting evidence, but GT...AG is not enough to identify introns reliably.

Most major-class eukaryotic introns begin with GT in DNA at the 5-prime donor and end with AG at the 3-prime acceptor. In RNA, the donor is written GU.

These two-base patterns occur frequently by chance. A long genomic sequence can contain thousands of possible GT...AG pairs that are never used for splicing.

Real splice recognition also depends on broader donor and acceptor context, a branch point, a polypyrimidine tract in many animals, RNA-binding proteins, transcription, and tissue-specific regulation.

The NCBI Bookshelf RNA-processing chapter describes the donor, acceptor, branch point, and consensus features used during pre-mRNA splicing.

GT-AG is a clue, not proof

Do not label every sequence between GT and AG as an intron. Confirm the interval with transcript alignment, curated annotation, RNA-seq junctions, or multiple compatible prediction signals.

What are the canonical splice-site sequences?

The canonical major-class intron usually has GT at its genomic 5-prime boundary and AG at its genomic 3-prime boundary when written in transcript orientation.

The exact consensus extends beyond those two bases. Boundary strength depends on neighboring positions, and many valid sites differ from the most common consensus.

Noncanonical introns exist. GC-AG introns are recognized by the major spliceosome, while a small minor-spliceosome class includes AT-AC and some other boundary combinations.

An NCBI sequence-analysis chapter notes that most splice sites follow GT-AG but lists biologically valid exceptions such as GC-AG and AT-AC. Read the overview.

Rejecting every non-GT donor can therefore remove real introns. Accepting every GT-AG pair creates the opposite error.

Method 1: Use a reference genome annotation

Reference annotation is the fastest method when the organism and genome assembly are known.

Search for the gene in Ensembl, NCBI Gene, or another authoritative organism database. Select the exact transcript and assembly, then inspect or download its exon coordinates.

For a plus-strand transcript with exons ordered from low to high genomic coordinate, each intron begins after one exon ends and ends before the next exon starts.

For 1-based inclusive exon coordinates:

exon 1: 101-180
exon 2: 301-420
exon 3: 501-620
intron 1: 181-300
intron 2: 421-500

The arithmetic is simple only after the coordinate convention is known. BED files normally use 0-based, half-open intervals, while GFF and many browser displays use 1-based, inclusive coordinates.

Ensembl’s transcript guide shows how exon and intron structure changes among splice variants. See the Ensembl training page.

How do you derive introns from exon coordinates in Python?

Sort exons in genomic order, then calculate the gap between every adjacent pair.

This example uses 1-based inclusive coordinates on one strand:

exons = [(101, 180), (301, 420), (501, 620)]
exons = sorted(exons)
introns = [
(left_end + 1, right_start - 1)
for (_, left_end), (right_start, _) in zip(exons, exons[1:])
]
print(introns)
# [(181, 300), (421, 500)]

For a minus-strand transcript, the genomic gaps are the same intervals, but transcript order runs from high to low coordinate.

Keep the strand attached to every interval. Sequence motifs must be read in transcript orientation, so a minus-strand intron requires the reverse complement before checking donor and acceptor bases.

Method 2: Align mRNA or cDNA to genomic DNA

Spliced alignment identifies introns by aligning mature transcript sequence across genomic gaps.

Use a high-quality mRNA, cDNA, or full-length transcript and the matching genomic assembly. A splice-aware aligner reports separate aligned blocks connected by introns.

NCBI Splign was designed to align transcript sequences to genomic DNA while detecting splice sites precisely. Read the NCBI Splign overview.

Other options include GMAP for transcript-to-genome alignment and minimap2 with splice-aware settings for long reads or transcript sequences.

A generic local alignment is often insufficient. It may split the transcript into fragments without placing exon boundaries at biologically plausible splice junctions.

A practical spliced-alignment workflow

A reliable workflow verifies sequence identity, strand, splice boundaries, and transcript completeness.

  1. Record the genome assembly and transcript accession.
  2. Remove vector or adapter contamination from the transcript sequence.
  3. Align the transcript to the correct genomic region or full assembly.
  4. Confirm one primary locus with high identity and coverage.
  5. Read aligned blocks as candidate exons.
  6. Read skipped genomic intervals as candidate introns.
  7. Inspect donor and acceptor motifs in transcript orientation.
  8. Compare the model with existing annotation and RNA-seq junction evidence.
  9. Report transcript-specific coordinates and the coordinate convention.

Full transcript coverage matters. A partial cDNA cannot define introns outside the sequenced portion, and sequencing errors near boundaries can shift alignments.

Method 3: Predict introns from genomic DNA alone

Genomic-only intron prediction combines splice motifs with coding potential, exon length, start and stop codons, conservation, and trained gene models.

Tools such as AUGUSTUS, GeneMark, and BRAKER infer complete gene structures rather than pairing every donor with every acceptor.

Prediction quality depends on species-specific training and evidence. Models trained on distant organisms may miss unusual exon structures or invent unsupported exons.

RNA-seq alignments, protein homology, and long-read transcripts can guide prediction. Evidence-supported annotation is generally more reliable than ab initio prediction alone.

Use genomic-only prediction as a hypothesis. Validate important splice junctions with transcript data, RT-PCR, or curated evidence.

How do RNA-seq reads show introns?

Splice-junction reads align across an exon-exon boundary and therefore support removal of the intervening genomic region.

A splice-aware RNA-seq aligner maps one part of a read to one exon and the remainder to a downstream exon. The skipped genomic interval is represented as an intron in the alignment.

In SAM or BAM format, the CIGAR operation N denotes a skipped reference region. For example, 50M1000N50M describes 50 aligned bases, a 1000-base genomic skip, and another 50 aligned bases.

One split read is weak evidence. Multiple uniquely mapped reads with consistent boundaries and appropriate strand support are more convincing.

Short-read RNA-seq may miss low-expression transcripts or confuse repetitive regions. Long-read RNA sequencing can connect several junctions within one molecule and help resolve isoforms.

What is the difference between an intron and an intergenic region?

An intron lies inside a specific precursor transcript, while an intergenic region lies between annotated genes or transcription units.

Both may be noncoding in a simple gene model, but their evidence and interpretation differ.

Introns are transcribed as part of pre-mRNA and removed by splicing. Intergenic regions may contain enhancers, noncoding transcripts, repeats, or unannotated genes.

Do not label a genomic gap as intronic unless two exons of the same transcript flank it.

How do alternative transcripts change introns?

Alternative splicing changes which genomic intervals are treated as exons or introns in the mature transcript.

Exon skipping joins nonadjacent exons. Alternative donor or acceptor use changes an intron boundary. Intron retention keeps an interval that is removed from another isoform.

Mutually exclusive exons and alternative first or last exons create additional transcript structures.

A single “gene sequence” is therefore not enough for precise intron reporting. Name the transcript version used for the analysis.

This matters for variant annotation. A change can be a splice-site mutation in one transcript but lie within an exon or untranslated region in another.

Common mistakes when identifying introns

The most common mistake is treating motif scanning as gene annotation.

Other frequent errors include:

  • Mixing genome assemblies
  • Using the wrong strand
  • Confusing coding sequence boundaries with exon boundaries
  • Mixing 0-based and 1-based coordinates
  • Ignoring transcript versions
  • Pairing exons from different isoforms
  • Treating every alignment gap as an intron
  • Ignoring noncanonical splice sites
  • Using protein-to-genome alignment without checking phase and splice evidence
  • Reporting an intron beyond the coverage of a partial transcript

Record the organism, assembly, transcript identifier, strand, coordinate system, alignment tool, and evidence source with every result.

How can you validate a predicted intron?

Validate a predicted intron with independent transcript evidence and consistent splice boundaries.

Strong support can include a curated transcript, several junction-spanning RNA-seq reads, a full-length cDNA, a long read spanning adjacent exons, or RT-PCR followed by sequencing.

Conservation in related species and a plausible splice motif add support but do not replace direct transcript evidence.

For experimental validation, design primers in the flanking exons. A cDNA product should be shorter than the corresponding genomic product by approximately the intron length.

Sequencing the cDNA amplicon confirms the exact exon-exon junction.

Key takeaways

The best way to identify introns is to use trusted transcript annotation or align mature RNA to genomic DNA with a splice-aware method.

Canonical GT...AG boundaries support a candidate but do not prove it. Valid noncanonical introns also exist.

Introns are transcript-specific. Alternative isoforms, strand orientation, genome assembly, and coordinate conventions must be reported.

Use genomic-only gene prediction when direct evidence is unavailable, then validate important junctions with RNA or experimental data.

Further reading

Read another related post

View all posts