Relative scalability of NoSQL databases for genotype data manipulation.

dc.contributorWAGNER ANTONIO ARBEX, CNPGL.
dc.creatorALMEIDA, A. L.
dc.creatorSCHETTINO, V. J.
dc.creatorBARBOSA, T. J. R.
dc.creatorFREITAS, P. F.
dc.creatorGUIMARÃES, P. G. S.
dc.creatorARBEX, W. A.
dc.date2018-12-26T23:42:22Z
dc.date2018-12-26T23:42:22Z
dc.date2018-12-26
dc.date2018
dc.date2018-12-26T23:42:22Z
dc.date.accessioned2026-07-07T03:47:18Z
dc.descriptionAbstract Genotype data manipulation is one of the greatest challenges in bioinformatics and genomics mainly because of high dimensionality and unbalancing characteristics. These peculiarities explains why Relational Database Management Systems (RDBMSs), the "de facto" standard storage solution, have not been presented as the best tools for this kind of data. However, Big Data has been pushing the development of modern database systems that might be able to overcome RDBMSs deficiencies. In this context, we extended our previous works on the evaluation of relative performance among NoSQLs engines from different families, adapting the schema design in order to achieve better performance based on its conclusions, thus being able to store more SNP markers for each individual. Using Yahoo! Cloud Serving Benchmark (YCSB) benchmark framework, we assessed each database system over hypothetical SNP sequences. Results indicate that although Tarantool has the best overall throughput, MongoDB is less impacted by the increase of SNP markers per individual.
dc.identifierRevista de Informática Teórica e Aplicada, v. 25, n. 2, p. 93-100, 2018.
dc.identifierhttp://www.alice.cnptia.embrapa.br/alice/handle/doc/1102528
dc.identifier10.22456/2175-2745.79334
dc.identifier.urihttp://hdl.handle.net/123456789/442901
dc.languageeng
dc.rightsopenAccess
dc.subjectDatabase
dc.subjectNoSQL
dc.subjectData Science
dc.subjectSNP
dc.subjectBioinformatics
dc.subjectGenotype
dc.titleRelative scalability of NoSQL databases for genotype data manipulation.
dc.typeArtigo de periódico

Archivos