Milo Yip has compared different itoa and dtoa implementations on Core i7, including my itoa algorithm 2, that use SSE2 instructions.
Results for itoa are interesting: SSE2 version is not as good as it seemed to be. Tricky branchlut algorithm is only 10% slower, moreover is perfectly portable. One obvious drawback of this method is using lookup-table - in real environment where is a big pressure on cache, memory access could be a bottleneck.
Pokazywanie postów oznaczonych etykietą sse2. Pokaż wszystkie posty
Pokazywanie postów oznaczonych etykietą sse2. Pokaż wszystkie posty
niedziela, 30 listopada 2014
sobota, 1 maja 2010
Speedup reversing table of bytes
With help of BSWAP instruction or SSE instructions (PSHUFD, PSHUFLW, PSHUFHW) or SSSE3 instruction (PSHUFB) reversing table can be faster. Speedup depends on three factors:
Read full article
- table size: larger=faster
- table address: aligned=faster/much faster (15.5 speedup - possible! see chart)
- CPU type
Read full article
Subskrybuj:
Posty (Atom)