(random unfinished engineering notes)

Wednesday, June 25, 2008

Profiling java apps using DTrace

- monitoring operational performance of java apps
- finding bottlenecks
- examining cache effects
- example : heavily IO-bound app
- OS/X / Solaris

idea :
(a) current java profiling options : (what/where)
(b) java options on system/level profiling (we can debug / but cannot to the strace-like detailed to-syscall-mapping, oprofile-like stuff etc)

dtrace - cool caps :



references :

http://www.devx.com/Java/Article/33943/0/page/3
http://www.solarisinternals.com/wiki/index.php/DTrace_Topics_Java
http://developers.sun.com/solaris/articles/java_on_solaris.html
http://developers.sun.com/solaris/articles/dtrace_ajax.html

Saturday, March 29, 2008

Managing shared state in Erlang

Though a functional language, with no apparent shared state, we can trivially implement state in erlang by (ab)using the single-bounded value concept (each value can be bound only once)

whereis(list_to_atom("portServer" ++ integer_to_list(PORT))).



Bloom filters for profit & fun

Perfect Hash Functions for fun and profit

Thursday, March 27, 2008

Breakpoint

Okie, cut here

This is another attempt into actually starting to write something ;)
So, stay tuned for a bunch of stupid posts :) I need something in order to get started. Hopefuly, I might be able to actually post something usefull some day. (in which case i will erase all the posts until that day :) ). So, some value after all .... these posts won't last forever .... :) :) read them while they are still here ... and have mercy :)

Monday, December 31, 2007

Linux buffer cache & how to disable it (and why ?)

Linux buffer cache provides a excelent mechanism for black-box performance optimization by a modest cost or max 2 memory hits (hit 1 -> page not found -> fetch buffer from disk to buffer cache (free heap space) -> hit 2 -> found in memory -> get block).

However, there are 2 cases when we would want to disable such behaviour :

(1) When 2-hit is too much of price to pay, we might want to think about direct io (O_DIRECT/ madvise()-style), and doing the memory buffer management by hand from userspace - often done for db cache management

(2) When we want to do unbiased benchmarking of heavily IO-dependent software (usually a single shot-benchmark is ok - pages are on the disk), but upont 2nd take , most of the pages remains in buffcache, which results in better performance, so no real metric can be imposed afterwards - so the only way to do the proper benchmark would be to disable buffacahe

Currently, we are interested in (2) , so here are couple of ideas how that could be done :

(a) create a non-trivial file of available memory size and write a simple code that mmap()'s it
(b) tune the swappiness kernel knob (/proc/sys/vm/swappiness) to 0 (proc memory over buffers), and fork() some ~64k dummy (nontrivial) processes :) - this should do the trick
(c) mounting the partition as raw device (no buffering then) - but this is usually highly impractical
(d) seting the O_DIRECT flag for every open() in the source (this is often tedious unless open() is invoked through a wrapper - a nice argument for doing so in such applications). We could write a simple wrapper (if possible for doing so):

int dopen(char *file) {
return open(file, O_DIRECT);
}

or in case of fopen() :

FILE *dfopen(char *file, char *pern) {
int fd = open(file, O_DIRECT | O_RDONLY);
return fdopen(fd,perm);
}

(note theat O_DIRECT is conditionaly defined by _GNU_SOURCE
, so don't forget to use -D_GNU_SOURCE flag when building)
(e) reboot the machine (if you are really desperate) :)
(f) allthough one might think that invoking "sync" from command line might do the trick - it actually just flushes *changed* blocks to the disk - which is not actually what we need (flushing the read-buffered blocks)
(g) the dirty way of doing (a) would be simply touching a very big file (for example dd if=/dev/urandom of=file_4GB bs=1024 count=4000000) and doing 'cat file_4GB > /dev/null'
(h) writing a simple code that allocates a huge chunk of memory and locks the allocated pages (using mlock() or similar) - thus reducing the available buffcache to a arbitrary size
(i) doing fcntl() with F_NOCACHE on all file descriptors in the source code - again quite tedious especially if there is no wrapper for open() call in the code
(g) using madvise() mechanism for telling kernel that the pages allocated won't be used in the future (which should result in kernel freeing the allocated pages from buffcache immidiatelly):

size_t dfread(void *ptr,size_t size,size_t nmemb, FILE *stream) {
size_t n;
n = fread(ptr,size,nmemb,stream);
madvise(ptr,size*nmemb,MADV_DONTNEED);
return n;
}


(h) if we're to just flush the entire buffercache, on the 2.6.16+ kernels we could use a a "drop caches" mechanism to free all pages from buffcache :
echo 1 > /proc/sys/vm/drop_caches

Sunday, December 03, 2006

Topics in compiler implementation

Some compiler-related ideas & problems (Appel : Modern compiler implementation in Java):
- Instruction selection : Machine instructions are represented by tree pattern. The optimal code is generated by selecting tree patterns which form the minimum tree. Thus, instruction selection becomes a task of tiling the tree with minimal set of tree patterms. There is a number of approaches to this problem, most notably, the Maximal Munch algorithm and Dynamic Programming.
In pratice, tree pattern - based approach can be a problem when optimizing for cics, rather than risc machine...
- Liveness Analysis : In order to optimize the register allocation, the compiler must analyze which variables are used ("alive") during every segment of program (for example, we can observe every basic block as one segment). Variables which are not used at the same time can be allocated in the same register. We observe the control-flow-graph of the program and determine the "live" variables at every edge of the graph. Then, we can perform graph partitioning in order to divide the program into segments with a number of alive variables which correspond to number of system registers. Again, we should do this in a optimal manner (every segment should allocate all registers, but among segments there should be minimum needed transfer of data between registers and memory).

Saturday, December 02, 2006

Notes on automata theory

Remarks from Motwani, Hopcroft & Ullman : Introduction to Automata Theory, Languages and Computation (during 45min reading) :
Chapter 2 : Finite Automata
- Give an example of practical communication/transaction protocol which can only be analyzed via automata model
- Examples of structural patterns which require formal verification of correctness ?
- Complexity bound for automata which require formal verification?
(even in the case of 2 simple, say 4-state automata, their single interconnection yields a product of automata resulting in 16-state automata )

- Optimal representation of DFA ? (space-efficient Transition Tables). Alternative methods of representation ?
- Algorithm-efficient representations ? (ex. finding all loops in automation by observing transition table -> can we speedup the algoritm by using structure other than table)

A NFA models the ability of input to "guess" the right state. Actually this can be seen as increasing complexity in order to simplify the description. The automata can reduce complexity of the problem by increasing the description. However, there are problems that cannot be treated in this manner.

A NFA is more descriptive than DFA, while there should be a 1-1 mapping NFA <-> DFA (Using a Tompson's algorithm , we can efficiently construct the DFA corresponding to NFA). Still, a NFA of size N corresponds to a DFA of size N+K . How do we prove the 1-1 mapping then ?

- [Offtopic] -> how can we efficiently create a program for automata generation via feeding it with a list of input examples and accept/reject results for each example ?

& stuff.....

Here it is... I will be typing some random stuff on development, cs & stuff...