Wednesday, August 7, 2013

Image processing using Data mining

Introduction to data mining images

Image processing is one of those things people are still much better at than computers.  Take this set of cats:



Just at a glance, you can easily tell the difference between the cartoon animals and the photographs.  You can tell that the hearts in the top left probably don’t belong, and that Odie is tackling Garfield in the top right.  The human brain does this really well on small datasets.

But what if we had thousands, millions, or even billions of images?  Could we make an image search engine, where I give it a picture of an animal and it says what type it is?  Could we make it automatically find patterns that people miss?

Yes!  This post is the beginning of a series about how.  Finding patterns in large databases of images is still an active research area, and these posts will hopefully make those results more accessible.  The current research still isn’t perfect, but it’s probably much better than you’d guess.

The “black box” framework
There are three basic steps in data mining images:
STEP 1: Create the “black box”
STEP 2: Cluster
STEP 3: Run queries
That’s it!
… well … sort of …

There are many different algorithms that can be used at each step.  Which ones you decide to use will depend on the type of information you’re mining from the images.  The rest of this post gives a high level overview of how each of these steps works, and later posts will focus on specific implementations for each step.

STEP 1: Creating the black box

The black box defines the “distance” between two images.  The smaller the distance, the more similar the images are.  For example:


Garfield is very similar to himself, that’s why Box A gives him a low score–nearly zero.  Odie is not very similar to Garfield, but he’s a lot closer than a palm tree.  The specific numbers outputted don’t matter.  All that matters is the ordering created by those numbers.  In this case:


Of course, if we compare against a different image, we will probably get a different ordering.
Likewise, we can get different orderings with a different black box.  Let’s imagine that Box A was designed to determine if two pictures are of the same type of animal.  If we test it on some new input, we might get:


Notice that Box A thinks the real cat is more similar to Garfield than Odie is.  Now let’s consider another black box.  Imagine Box B is designed to see if two images were drawn in a similar style.  Box B might give the following:


Now, Odie is similar to Garfield (they’re both drawn by Jim Davis), but the cat is no longer similar to Garfield (because it’s a photograph).  Box B gives the opposite results of box A.
Creating a good black box is the hardest part of data mining images.  Most research is dedicated to this area, and most of this series will be focused on evaluating the performance of different black boxes.  Which ones are good depends on your dataset and what information you’re trying to extract.  Some general categories of black boxes we’ll look at are:
  1. Histogram analysis (a simple technique that can be surprisingly effective on colored input)
  2. Converting images into a time series (for analyzing the shapes of rigid objects, e.g. fruit)
  3. Creating shock graphs (for analyzing the shapes of non-rigid objects, e.g. animals)
  4. Komolgorov comlexity of the images (for comparing an image’s textures)
But first, let’s take a closer look at what makes a black box good.

Properties of a good black box


There are two more aspects of black boxes we must look at.  First, every black box will be sensitive to certain features of an image and invariant to others.  In the examples below, Box C is sensitive to shape, but invariant to color.  Box D is sensitive to color, but invariant to shape.




Most black box algorithms contain both sensitivities and invariances.  These are the properties you will use to decide which black box is best for your application.
Second, a black box is a metric and as such must satisfy four criteria:
  1. distance(x, y) ≥ 0     (non-negativity)
  2. distance(x, y) = 0   if and only if   x = y     (identity of indiscernibles)
  3. distance(x, y) = distance(y, x)     (symmetry)
  4. distance(x, z) ≤ distance(x, y) + distance(y, z)     (subadditivity / triangle inequality).
If you don’t understand these criteria, don’t worry too much.  All the black boxes we’ll look at in the rest of this series will satisfy these criteria automatically.

STEP 2: Cluster the images

Clustering is much easier than designing the black box.  Clustering algorithms are used in many fields, so they have received much more attention.  Some clustering algorithms commonly used are:
  1. Support vector machines
  2. K-means
  3. Neural networks
  4. Hierarchical clustering (i.e. Dendrograms)
There are many more as well.  In general, you can use whatever clustering algorithm you want.  When developing an application, most people will try several and pick whichever one happens to work the best for their data.

Here’s an example clustering of our cat data using Black Box A (i.e. by what the picture is of):



We’ve created three clusters.  The red cluster contains hearts, the white cluster contains cats, and the blue cluster is an anomaly.  It contains both a cat and a dog, and there is no easy way to separate them.  If we had used a hierarchical classifier, the “contains cats and dogs cluster” might be a sub-cluster of the “contains cats cluster.”

Here’s the same data clustered using Black Box B (i.e. by how the picture is drawn):



Now we have only two clusters: the white cluster contains cartoons, and the red cluster contains photographs.One last note.  Most of the CPU work gets done during this step.  On large datasets, clustering can take hours to months depending on the algorithm and the speed of the black box.  There are many tricks for speeding up clustering, which will take a look at in later posts.


STEP 3: Run your query

Queries are fairly easy once the ground work is set up with the black box and clustering. Sometimes, all you want to know is how STEP 2 clustered your input.  For example, you could query “how many types of animals are in this dataset?”  The answer would just be the number of clusters using Box A.  Typically, however, your query we will supply the database with an image and find similar images.


If we’ve done steps 1 and 2 well, this should take only seconds even when the database contains millions of images.  Of course, it’s not always possible to do steps 1 and 2 well enough to make this happen.  Later posts may cover some new techniques for speeding up the querying process.

The Rest of the Series

So far, we’ve seen that the black box framework for image data mining is very simple:

STEP 1: Create the “black box”
STEP 2: Cluster
STEP 3: Run queries

The tricky part is putting the right algorithm in each step.  In the rest of the series, we’ll look at a few different black boxes, and show how to efficiently combine them with a clustering algorithm.  The different types of black boxes are the most interesting part of image mining, so we will focus on that first.

 

Tuesday, July 9, 2013

Basic Steps in Ubuntu 12.10


CLASSPATH SETTING IN UBUNTU

sudo gedit etc/environment

enter your password

PATH=".:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games" (already exist)
just add below two lines
 

here i am using java6

JAVA_HOME="/usr/lib/jvm/
java-6-openjdk"

CLASSPATH=".:/usr/lib/jvm/java-6-openjdk/lib"


 Disable Guest Session and the Remote login in UBUNTU


To disable the guest session and/or remote logon in Ubuntu 12.10 (Quantal Quetzal):
  1. Open a terminal window.
  2. Type “sudo nano /etc/lightdm/lightdm.conf”
  3. Type in your sudo password.
  4. Add the following in a new line at the end of the file if you want to disable the guest session: allow-guest=false
  5. Add the following in a new line at the end of the file if you want to disable the remote login option: greeter-show-remote-login=false
  6. If you choose to disable both, your lightdm.conf file should probably look like this:
    [SeatDefaults]
    greeter-session=unity-greeter
    user-session=ubuntu
    allow-guest=false
    greeter-show-remote-login=false
  7. Hit CTRL-X to exit nano.
  8. Hit Y to save the file.
  9. Hit Enter to accept the original filename and overwrite the file.
  10. On the terminal, type “sudo restart lightdm” to restart the display manager. Doing so will require you to log in again and may close running apps, so save your work before doing so.
  11. The guest session and/or remote logon options should now be disabled.

UNIX Basic Commands

Files

  • ls --- lists your files
    ls -l --- lists your files in 'long format', which contains lots of useful information, e.g. the exact size of the file, who owns the file and who has the right to look at it, and when it was last modified.
    ls -a --- lists all files, including the ones whose filenames begin in a dot, which you do not always want to see.
    There are many more options, for example to list files by size, by date, recursively etc.
  • more filename --- shows the first part of a file, just as much as will fit on one screen. Just hit the space bar to see more or q to quit. You can use /pattern to search for a pattern.
  • emacs filename --- is an editor that lets you create and edit a file. 
  • mv filename1 filename2 --- moves a file (i.e. gives it a different name, or moves it into a different directory (see below)
  • cp filename1 filename2 --- copies a file
  • rm filename --- removes a file. It is wise to use the option rm -i, which will ask you for confirmation before actually deleting anything. You can make this your default by making an alias in your .cshrc file.
  • diff filename1 filename2 --- compares files, and shows where they differ
  • wc filename --- tells you how many lines, words, and characters there are in a file
  • chmod options filename --- lets you change the read, write, and execute permissions on your files. The default is that only you can look at them and change them, but you may sometimes want to change these permissions. For example, chmod o+r filename will make the file readable for everyone, and chmod o-rfilename will make it unreadable for others again. Note that for someone to be able to actually look at the file the directories it is in need to be at least executable.
  • File Compression
    • gzip filename --- compresses files, so that they take up much less space. Usually text files compress to about half their original size, but it depends very much on the size of the file and the nature of the contents. There are other tools for this purpose, too (e.g. compress), but gzip usually gives the highest compression rate. Gzip produces files with the ending '.gz' appended to the original filename.
    • gunzip filename --- uncompresses files compressed by gzip.
    • gzcat filename --- lets you look at a gzipped file without actually having to gunzip it (same as gunzip -c). You can even print it directly, using gzcatfilename | lpr
  • printing
    • lpr filename --- print. Use the -P option to specify the printer name if you want to use a printer other than your default printer. For example, if you want to print double-sided, use 'lpr -Pvalkyr-d', or if you're at CSLI, you may want to use 'lpr -Pcord115-d'. See 'help printers' for more information about printers and their locations.
    • lpq --- check out the printer queue, e.g. to get the number needed for removal, or to see how many other files will be printed before yours will come out
    • lprm jobnumber --- remove something from the printer queue. You can find the job number by using lpq. Theoretically you also have to specify a printer name, but this isn't necessary as long as you use your default printer in the department.
    • genscript --- converts plain text files into postscript for printing, and gives you some options for formatting. Consider making an alias like alias ecop 'genscript -2 -r \!* | lpr -h -Pvalkyr' to print two pages on one piece of paper.
    • dvips filename --- print .dvi files (i.e. files produced by LaTeX). You can use dviselect to print only selected pages.

Directories

Directories, like folders on a Macintosh, are used to group files together in a hierarchical structure.
  • mkdir dirname --- make a new directory
  • cd dirname --- change directory. You basically 'go' to another directory, and you will see the files in that directory when you do 'ls'. You always start out in your 'home directory', and you can get back there by typing 'cd' without arguments. 'cd ..' will get you one level up from your current position. You don't have to walk along step by step - you can make big leaps or avoid walking around by specifying pathnames.
  • pwd --- tells you where you currently are.

Finding things

  • ff --- find files anywhere on the system. This can be extremely useful if you've forgotten in which directory you put a file, but do remember the name. In fact, if you use ff -p you don't even need the full name, just the beginning. This can also be useful for finding other things on the system, e.g. documentation.
  • grep string filename(s) --- looks for the string in the files. This can be useful a lot of purposes, e.g. finding the right file among many, figuring out which is the right version of something, and even doing serious corpus work. grep comes in several varieties (grepegrep, and fgrep) and has a lot of very flexible options. Check out the man pages if this sounds good to you.