Conversion numbers to hexadecimal representation - SWAR, plain SSE, and draft of BMI2 implementation.
Article SSSE3: printing hex values describes the same topic but is limited to exploit PSHUFB.
niedziela, 21 września 2014
czwartek, 18 września 2014
String literals are weird (at least in C/C++)
Simple quiz: what is the length of this string "\xbadcafe"?
- 5 letters
- 4 letters
- 2 letters
- 1 letter
czwartek, 11 września 2014
Python: rename file in Windows
Observation: function os.remove under Windows doesn't allow to overwrite an existing file. Python from Cygwin works properly.
Here is a workaround:
Here is a workaround:
def rename_file(old, new):
if sys.platform != 'win32':
os.rename(new, old)
else:
os.remove(old)
os.rename(new, old)
Conversion numbers to binary representation
New article Conversion numbers to binary representation — SIMD & SWAR versions.
Few years ago I've described an MMX variant of SIMD algorithm. But the text was in Polish, so audience was limited.
Few years ago I've described an MMX variant of SIMD algorithm. But the text was in Polish, so audience was limited.
sobota, 22 marca 2014
C++ bitset vs array
C++ bitset conserves a memory, but at cost of speed access. Bitset must be slower than set represented as a plain old array, at least when sets are small (say few hundred elements).
Lets look at this simple functions:
The file was compiled with g++ -std=c++11 -O3 set_test.cpp; Assembly code of the core of any_in_byteset:
Now, look at assembly code of any_in_bitset:
All these instructions implements the if statement! Again we have load from memory (5f), but checking which bit is set require much more work. Input (edx) is split to lower part --- i.e. bit number (67, 6c) and higher part --- i.e. word index (6c). Last step is to check if a bit is set in a word --- GCC used variable shift left (6f), but x86 has BT instruction, so in a perfect code we would have 2 instructions less.
However, as we see simple access in the bitset is much more complicated than simple memory fetch from byteset. For short sets memory fetches are well cached and smaller number of instruction improves performance. For really large set cache misses would kill performance, then bitset is much better choice.
Lets look at this simple functions:
// set_test.cpp
#include <stdint.h>
#include <bitset>
const int size = 128;
typedef uint8_t byte_set[size];
bool any_in_byteset(uint8_t* data, size_t size, byte_set set) {
for (auto i=0u; i < size; i++)
if (set[data[i]])
return true;
return false;
}
typedef std::bitset<size> bit_set;
bool any_in_bitset(uint8_t* data, size_t size, bit_set set) {
for (auto i=0u; i < size; i++)
if (set[data[i]])
return true;
return false;
}
The file was compiled with g++ -std=c++11 -O3 set_test.cpp; Assembly code of the core of any_in_byteset:
28: 0f b6 10 movzbl (%eax),%edx 2b: 83 c0 01 add $0x1,%eax 2e: 80 3c 11 00 cmpb $0x0,(%ecx,%edx,1) 32: 75 0c jne 40 34: 39 d8 cmp %ebx,%eax 36: 75 f0 jne 28Statement if (set[data[i]]) return true are lines 28, 2e and 32, i.e.: load from memory, compare and jump. Instructions 2b, 34 and 36 handles the for loop.
Now, look at assembly code of any_in_bitset:
5f: 0f b6 13 movzbl (%ebx),%edx 62: b8 01 00 00 00 mov $0x1,%eax 67: 89 d1 mov %edx,%ecx 69: 83 e1 1f and $0x1f,%ecx 6c: c1 ea 05 shr $0x5,%edx 6f: d3 e0 shl %cl,%eax 71: 85 44 94 18 test %eax,0x18(%esp,%edx,4) 75: 75 39 jne b0
All these instructions implements the if statement! Again we have load from memory (5f), but checking which bit is set require much more work. Input (edx) is split to lower part --- i.e. bit number (67, 6c) and higher part --- i.e. word index (6c). Last step is to check if a bit is set in a word --- GCC used variable shift left (6f), but x86 has BT instruction, so in a perfect code we would have 2 instructions less.
However, as we see simple access in the bitset is much more complicated than simple memory fetch from byteset. For short sets memory fetches are well cached and smaller number of instruction improves performance. For really large set cache misses would kill performance, then bitset is much better choice.
środa, 19 marca 2014
Is const-correctness paranoia good?
Yes, definitely. Lets see this simple example:
$ cat test.cpp
int test(int x) {
if (x = 1)
return 42;
else
return 0;
}
$ g++ -c test.cpp
$ g++ -c -Wall test.cpp
int test(int x) {
if (x = 1)
return 42;
else
return 0;
}
Only when we turn on warnings, compiler tell us about a possible error.
Making the parameter const shows us error:
$ cat test2.cpp
int test(int x) {
if (x = 1)
return 42;
else
return 0;
}
$ g++ -c test.cpp
test2.cpp: In function ‘int test(int)’:
test2.cpp:2:8: error: assignment of read-only parameter ‘x’
if (x = 1)
^
All input parameters should be const, all write-once variables serving as a parameters for some computations should be also const.
Quick and dirty ad-hoc git hosting
Recently I needed to synchronize my local repository with a remote machine, just for full backup. It's really simple if you have standard Linux tools (Cygwin works too, of course).
1. in a working directory run:
2. in a parent directory start HTTP server:
3. on a remote machine clone/pull/whatever:
1. in a working directory run:
$ pwd /home/foo/project $ git update-server-info
2. in a parent directory start HTTP server:
$ cd .. $ pwd /home/foo $ python -m SimpleHTTPServer Serving HTTP on 0.0.0.0 port 8000 ...
3. on a remote machine clone/pull/whatever:
$ git clone http://your_ip:8000/project/.gitStep 1 have to be executed manually when local repository has changed.
Subskrybuj:
Posty (Atom)