Monday, December 7, 2015

How can I use docker to containerize my data analytics app: A general overview

I recently answered this question on Quora about Dockerizing data analytics application, intriguing a thought on starting my series of posts on Docker.

This is my first post on that series. As opposed to traditional ways of a series, I will touch very little of preliminary details on Docker and it's purpose. I will take a straight dip into design and architecture aspects of dockerizing a system.

First thing first
Docker is a very light-weight application engine that deploys VM-like containers that shares system level resources to allow easy deploy and multi-tenancy. But it has its own network and process space, as well as a layered union mount file system. It's all written in Go. The three main components are:
·        docker client,
·        docker daemon or server (REST API)
·        docker containers.

How big is the docker container?
The docker container rootfs (lossely speaking, the operating system FS layer) and tmpfs can be of any size depending on the service you are dockerinzing. A small python app can be under 1MB or an full blown service can range around 16G or whatever it is configured to be.
The docker image can be a few hundred megabytes, if that is what you are asking. Usually it takes fractions of a second to launch a container from an image.
Can I put a data analytics product in it
Like I said, yes, you can. But let's break the problem down here.
·        Services - Docker centers around Service Oriented Architecture, or SOA. How will you you reorganize your application into micro-level, self-sufficient services that can communicate with each other? Let's say you have a web app. You need the web engine server (WARs) to be dockerized and there are plenty of examples on the internet to do this. You sure have a database instance, and that can be in a container. Then say you have a few daemons running for something - you need to make a call on where to put them. In short,  the key design principle is to identify the services to dockerize. Then maybe start with writing your own dockerfile for one component and get the ball rolling from there.
·        Networking - Docker solves the port-conflicts in multi-tenancy by dynamic mapping of ports. Each docker container has configurable and statically mapped ports exposed to the user that maps to physical ports in the system (a process abstracted by docker). Docker containers also have IPs assigned that are not discoverable outside the host. In case of service colocation being absent, you might also need the host IPs or configure the docker containers with unique discoverable IPs.
·        The data - Docker does not work with all the filesystems, so based on how your files and other data is stored, it might become a tall order. But in general, you can expose a volume on the host, or even dockerize a volume, and make it available to dockerized services. So yes, data can be ported too.
·        Handling - You might want a resource manager like YARN to allocate container. Zookeeper or Consul can take care of failover. Consul has built in support for configuration management too.


Monday, April 6, 2015

Unpaired element in array

I started using codility after long, and turned to a very basic problem. It was easy and did not take me more than 10 minutes to solve. I worked out the testcases by hand, and then was browsing a bit on Java basics after I finished the code, then submitted it. Otherwise it was so easy that I doubt the possibility of this occurring on any interview.

Still I felt that it was beneficial to solve it because:
>> It helped me jargon bitwise operators in mind ensuring I wouldn't mistake the syntaxes due to rustiness (we bearly use bitwise in day-to-day job, do we?)
>> It gave me a confidence boost to crack the best solution with perfect testcases / edge cases in one go.

Problem:
A non-empty zero-indexed array A consisting of N integers is given. The array contains an odd number of elements, and each element of the array can be paired with another element that has the same value, except for one element that is left unpaired.
For example, in array A such that:
  A[0] = 9  A[1] = 3  A[2] = 9
  A[3] = 3  A[4] = 9  A[5] = 7
  A[6] = 9
  • the elements at indexes 0 and 2 have value 9,
  • the elements at indexes 1 and 3 have value 3,
  • the elements at indexes 4 and 6 have value 9,
  • the element at index 5 has value 7 and is unpaired.
Write a function:
class Solution { public int solution(int[] A); }
that, given an array A consisting of N integers fulfilling the above conditions, returns the value of the unpaired element.
For example, given array A such that:
  A[0] = 9  A[1] = 3  A[2] = 9
  A[3] = 3  A[4] = 9  A[5] = 7
  A[6] = 9
the function should return 7, as explained in the example above.
Assume that:
  • N is an odd integer within the range [1..1,000,000];
  • each element of array A is an integer within the range [1..1,000,000,000];
  • all but one of the values in A occur an even number of times.
Complexity:
  • expected worst-case time complexity is O(N);
  • expected worst-case space complexity is O(1), beyond input storage (not counting the storage required for input arguments).
Elements of input arrays can be modified.

My solution that received 100%
class Solution { public int solution(int[] A) { int temp = 0; for(int i = 0; i < A.length; i++) { temp = temp ^ A[i]; } return temp; } }

Saturday, March 1, 2014

Maximum Binary Gap : O(log n) solution

A binary gap within a positive integer N is any maximal sequence of consecutive zeros that is surrounded by ones at both ends in the binary representation of N.
For example, number 9 has binary representation 1001 and contains a binary gap of length 2. The number 529 has binary representation1000010001) and contains two binary gaps: one of length 4 and one of length 3. The number 20 has binary representation 10100 and contains one binary gap of length 1. The number 15 has binary representation 1111 and has no binary gaps.
Write a function:
class Solution { public int solution(int N); }
that, given a positive integer N, returns the length of its longest binary gap. The function should return 0 if N doesn't contain a binary gap.
For example, given N = 1041 the function should return 5, because N has binary representation 10000010001 and so its longest binary gap is of length 5.
Assume that:
·    N is an integer within the range [1..2,147,483,647].
Complexity:
expected worst-case time complexity is O(log(N));
expected worst-case space complexity is O(1).

// My perfect score solution
class Solution {
    public int solution(int N) {
        if(N < 1)
            return -1;    
        int res = 0;
        int gapLen = 0;
        boolean binGapStart = false;
        boolean gapZeroes = false;           
        while(N >= 1) {
            if(N%2 == 1) {
                binGapStart = binGapStart && gapZeroes;              
                if(binGapStart) {
                    res = (res > gapLen)? res:gapLen;
                    gapZeroes = false;
                    gapLen = 0;
                }
                else
                    binGapStart = true;
            }
            else{
                gapZeroes = binGapStart;
                if(gapZeroes)
                    gapLen++;
            }           
            N = N/2;
        }       
        return res;       
    }
}


Java Tidbidz 1: Access modifiers, Arrays

  • Access Modifiers - All members of interfaces are implicitly public. It is, in fact, a compile-time error to specify any access specifier for an interface member other than public (although no access specifier at all defaults to public access).
  • Array - Arrays are special objects in java with no "class definition" (no .class file).
  • Array.length [public final int variable], but String.length()

Java tidbidz 2: Abstract class and Interface

  • Interface variables static and final by default [interfaces cannot be instantiated in their own right; the value of the variable must be assigned in a static context in which no instance exists. The final modifier ensures the value assigned to the interface variable is a true constant that cannot be re-assigned by program code.]
  • We can have abstract class without abstract method but not abstract method in non-abstract class, because declaring a class abstract only means that you don't allow it to be instantiated on its own, while an abstract method must be defined by subclasses.
  • You can even have abstract classes with final methods but never final classes with abstract methods.

Saturday, January 11, 2014

Detect and identify cycle in linked list

Tortoise-Hare Algorithm illustrated -

Detection:
The classic Tortoise-Hare algorithm (Floyd Cycle Detection) requires two pointers, one fast (Hare or H) and a slower (Tortoise or T). Start both of them from the same (head) node and run the Hare twice as fast as the Tortoise. Clearly, they will meet at some point if there is a cycle. So far so good, very clear and intuitive.

Identification:
Let's assume that the singly linked list is represented as x = {x(0), x(1), ..., x(z)}. It has a cycle of length n starting at node x(m). T and H meet at node x(k), which is i distance away from start of the cycle. [For simplicity, I am eliminating obvious assumptions like, i <=n, k>=m, etc).



As H is twice as fast as T, by the time T crosses k nodes, H will traverse 2k nodes. Hence, we can write the following equations,

For T: k = m + i
For H: 2k = m + n + i
=> 2(m + i) = m + n + i
=> m =n - i

Based on this observation, T is restarted from the head or x(0), and both T and H are moved at the same speed of 1 node per time.
You will see that they ALWAYS meet at start of the cycle. This is because, by the time T crosses m nodes to reach x(m), H will now also cross m nodes. But H is moving in the cycle. So, it will cross, remaining (n-i) nodes of cycle (as it was at node x(k), where k = m + i) and reach the beginning of the cycle, x(m). But hey! (n-i) is actually m! So, H has just crossed m nodes, and will stop at x(m), where T has just reached :)
Thus, we know where the cycle starts.
The reason I wrote this blog is because the usual descriptions of this algorithm just tells you to restart T from head and move them at the same speed without this little description (which can be intuitive for many). But I felt the need of explaining the reason.

The rest is simple. Keep either T or H fixed at x(m), move the other till it reaches the fixed pointer again. The number of nodes traversed is length of the cycle.

Tuesday, February 7, 2012

Realizing that it's time for the dissertation dread

Well, it is this time of the life again. I am thinking of graduation, and thinking very hard. So hard that I can say "I am graduating this year". So here comes the most crucial phase of my PhD - the dissertation.

I have started working on my PhD thesis (I like calling it thesis too) sometime back, but never did I realize it so evidently and obviously that it is that very period of my PhD. It is the final stage, and I have read so much about it already. Starting from PhD Tips, Time Management for Dissertation, Secrets of Successful PhD Students, What They Don't Tell You, How to Write a Good PhD Dissertation, I never missed PhD Comics, GradCafe, How to Improve your Concentration, Resume and Thesis-writing workshops. I attended seminars not only in my own department, but in other departments and watched research-talk videos online - all keeping my overall research growth in mind, which is supposed to converge concisely in my thesis. I devoted considerable amount of time, thought, anxiety and finally panic on my dissertation. Now I am in a stage of actually working on the dissertation dread and getting the actual thesis out of the mess.

Yes! That's a very important thing. My thesis is there, somewhere at the back of my mind, hear soul, knowledge, library, internet and ofcourse, my supervisor. I now realized that I need to cultivate it.

I'm in this phase where I am writing publications parallel to my thesis. I am reading a lot and thinking even in my shower or while driving (well, not that I suggest that to anyone). I will be posting more often here, as I ride this tide.

Tuesday, June 21, 2011

Analyzed the data files that I've mentioned about in the previous post. Turns out that Machine Learning is a novel and effective choice for Sybil attack detection in VANETs. Only 10% of the nodes with varying support levels were varied. Details of this work will be soon published as a part of my VANET security survey journal and probably be discussed also in the Sybil Attack Detection Conference paper. Was a good learning experience with Weka for about a week.
Our basic assumption was that the nodes use up all the possible fake IDs it has. For example, if you are a moving vehicle with 10 IDs V1....V10 (and your original ID is V0) that you have fabricated, you are posing as V0, V1,....V10 simultaneously. So if we keep confidence level 1 and perform the analysis, we were able to detect all Sybil nodes without any false positives!!! yayyyy!!!!! that's so exciting and we had our day - but next morning (in this case a couple of weeks later) after couple of cups of coffee (read couple of other brain-storms on independent research issues) we figured out, that's not-only-dumb-but-meaningless assumption. Why would anybody use up all the aces (fake IDs) and shout out loud "look I'm the mischiefer.....catch meeee".....So we started figuring out ways to deal with probabilistic distribution a malicious node might follow to use up fake IDs. That's my Friday night companion tonight........let's see......

Monday, June 6, 2011

I got a bunch of data files couple of days back. The data is collected from a simulator processed with real traces. Details of the simulator will be updated shortly after I get to talk to my fellow labmate whose project has been assigned to me. I am pretty excited with the data as it looks pretty huge and I am not a data-mining person. 
The data is about vehicle traces. Vehicles' connections and IDs are listed over time instants. I need to now analyze this data to find the outlier vehicles faking IDs which will help detect sybil attack. Reading from platoon dispersion in urban areas, this sounds like an interesting technique. Will post the progress and results time-to-time once I get started with the analysis on Weka.