Classification of nucleotide sequences by latent semantic analysis

DNA sequences are analyzed using latent semantic analysis. A set of nucleotide sequences is received in which the set has a first number of sequences. A set of basis vectors is determined, in which the set has a second number of basis vectors, the second number being smaller than the first number. E...

Full description

Saved in:
Bibliographic Details
Main Authors Sayood Khalid, Way Sam, Garrity George, Nalbantoglu Ozkan Ufuk
Format Patent
LanguageEnglish
Published 23.05.2017
Subjects
Online AccessGet full text

Cover

Loading…
More Information
Summary:DNA sequences are analyzed using latent semantic analysis. A set of nucleotide sequences is received in which the set has a first number of sequences. A set of basis vectors is determined, in which the set has a second number of basis vectors, the second number being smaller than the first number. Each basis vector represents a specific combination of predetermined nucleotide segments. For each of the nucleotide sequences, an approximate representation of the nucleotide sequence is determined based on a combination of the basis vectors. For each pair of nucleotide sequences, a distance between the pair of nucleotide sequences is determined according the distance between the approximate representation of the pair of nucleotide sequences. The set of nucleotide sequences are classified based on the distances between the pairs of nucleotide sequences.
Bibliography:Application Number: US201313954925