Category: programming

  • Python vs Matlab/Octave: anonymous functions

    I learned Matlab before I learned Python, and some of the patterns one learns in one language are hard to forget when you learn a new one. Scoping is a common area of misunderstanding between the two languages.

    I’m using Octave in these examples, but the language is the same in this respect. Consider the following snipped of code:

    > a = 1;
    > b = 2;
    > f = @(x) a*x + b;
    > f(2)
    ans =  4
    > a = 0;
    > b = 0;
    > f(2)
    ans =  4
    

    What we see here is that Matlab/Octave binds the values of the variables to the function defined in line 3 at the time of definition. This means changing a and b later doesn’t change f. This is useful when functions are passed back and forth – you can build a function in one place and use it later without fearing that changing local variables will change the behaviour of the function. Let’s do the same thing in Python (the function definition below is completely equivalent to f = lambda x: a*x + b):

    a = 1
    b = 2
    def f(x):
        return a*x + b
    print f(2)
    a = 0
    b = 0
    print f(2)
    

    Output:

    4
    0
    

    If you are used to the Matlab way, this result may be counter-intuitive, but I remember teaching students to program in Matlab, and in fact this was a common problem. Remembering that the variables are captured in Matlab was hard for some people, who expected Python’s behaviour.

    It is worth noticing that in the Python version, the names bound to f are not simply evaluated in whatever scope they are passed to, as can be seen if we define a new function (which will have its own namespace):

    def g():
        a = 5
        b = 6
        return f(2)
    print g()
    

    Output:

    0

    The rule in Python is therefore that the names in functions are resolved when the function was defined, but that the values will only be obtained at the time of execution of the function. Because g has its own namespace, that a and b are different from the one in the global namespace.

    In fact, it is possible to think of what is happening in a similar way if you remember how names and values work in Python. I highly recommend this article for a good overview of how this works.  I have also found this Online Python Tutor very helpful to visualise how the name-value relationship works.

  • The impact of environment: how much do IDEs contribute?

    Those who know me will know this is a much-beloved topic of mine. I am setting this stuff down in an attempt to purge some of these thoughts from my brain.

    During a conversation with a colleague yesterday, the idea of language productivity came up. It’s an old argument: sure, you can write less lines of code in say Python than in Java (see my post on this topic), but with a modern IDE, you’ll end up ahead as the IDE will be doing a lot of that typing for you. This forms part of a bigger conversation about the relative power of different languages, the relative merits of various IDEs and the way the brain works, before and after exposure to various models of computation. So, I’ve been maintaining a variety of links to research on these topics, and I’ll try to state my current understanding succinctly in the next couple of paragraphs. I would really value any feedback that you could give me (for the handful of people who actually read this). Although I enjoy reading anecdotes of personal experiences, I would appreciate most if you could point to proper scientific studies on the topic.


    So, this study, from 2000, compared C, C++, Java, Perl, Python, Rexx, and Tcl for a particular programming task: in this case a program to find which words could be spelled with particular phone numbers using the normal keypad letter mappings. They found that the “scripting language” programs were significantly shorter and took less time to write, but did not perform significantly poorer than the compiled languages. An interesting feature is that the compiled languages (C, C++ and Java) required considerably longer programs and took longer to write. The lines/hour for all languages were roughly comparable. They don’t mention which IDEs the programmers were using, but one assumes that at least some of them were using “modern” IDEs. A 2003 interview with Guido van Rossum, highlights this idea of Python requiring less “finger typing”. He’s obviously biased as the creator of Python, but considering that he also wrote the Python interpreter (in C) and has a large realm of experience in various languages, one would think he has had some time to experience the differences.
    Another study, from 2009 may explain a part of this. They found that reading code was largely a function of the number of tokens. This is important, because it implies that less tokens can aid in understanding already written code and have an effect on writing it. Importantly, they also find that simple domain mappings are easier to comprehend than complicated ones.
    I have not been able to find much quantitative research on the benefits of IDEs. Although I acknowledge that environment makes a big difference, I tend to think about code in my head without an IDE, which makes the actual language quite important for me. To some extent one must acknowledge that there is a personal aspect as well – some people prefer a particular kind of approach, as shown in this article about language mavens vs tool mavens, which I’ve linked to before. I’ve slowly been learning Eclipse, and it is quite amazing how much of the mundane Java stuff it can automate. I’m also a great believer in refactoring tools (I use rope, with bindings for Emacs). At the same time it is clear to me that much of the stuff you get in Eclipse could be addressed by streamlining the language a bit.
    Finally, the larger picture about fitting into an environment. This is where the established languages have the upper hand. Popular languages are more widely known, have more books and environments available and have the corporate OK. Also, if the libraries of the operating system you are using have been crafted for a particular language, there is a clear advantage to aligning yourself with that. As a researcher, though, it is often a bit of a conundrum: the real publishable results are more important than the interface, so you find the majority of computational research code has no GUI, and is only lightly coupled to the host OS. This is especially true of parallel code which inevitably has to run on headless cluster boxes.
    I don’t want to make this another post about choosing a programming language, but I will restate my final strategy: use the most expressive language available for prototyping – this will inevitably be a dynamically typed language with a REPL. For me, this is Python. This hooks into the constant lines per hour result of before. When the program produces the correct output and the algorithm is set, re-implement the slow bits in a lower level language. For me this is Fortran, but it could be C++ or whatever has good support for your problem domain. This strategy has placed me a in a poor position to benefit from IDE support. For that, you typically want to do everything in a combination of IDE and well-supported language (this is the classic Java/C++/C# for everything approach).
    Anyone with links to proper quantifiable research on programmer productivity (especially the effect of IDEs)?
  • The fallacy of the general purpose tool

    I carry a Leatherman on my belt every day wherever I go. It has become a bit of an extension of my body, which I only really notice when I have to fly somewhere and I’m forced to take it off for the flight. Then I get to my destination and I can’t open the cable ties on my luggage. At the same time, I have a whole pinboard of tools up in my garage, which overlap somewhat with my Leatherman: a dedicated set of pliers, some dedicated screwdrivers and scissors and so on. This is because the Leatherman is not as good as any of these dedicated tools.

    I think the same goes for programming languages, and therefore I have learned quite a few, just like I have stocked my pinboard with tools that do one thing very well. Many people find learning new languages onerous and will try to find the one language that does everything they want to do well. Their arguments against learning a new language often involves something like “but I can do that just fine in my language, and then I don’t have to learn a new language”. Here’s an example from real life that illustrates the advantages of special purpose languages.

    My father approaches me with this problem. He uses an accounting package which can import OFX files (they’re an SGML/XML format for financial records), but there is a slight problem. The ZAR needs to be changed to USD and the dates (which are tags of the form <DT20111201> – yes, no closing tag as it’s SGML) have to be changed from YYYYMMDD to YYYYDDMM. Now, turns out that Python has a nice OFX module. So does Java. But the fastest way to make the changes to his files is probably sed:

    sed -r -i -e 's/ZAR/USD/' -e 's/<DT(.*)><dt></dt><dt>([0-9]{4})([0-9]{2})([0-9]{2})/<DT\1>\2\4\3/' filename.ocx</dt>

    This does the replacement in-place, is really fast and is quick enough to throw together. This is the kind of thing that sed shines at. In fact, it is exactly what it was designed to do, so it is unsurprising that it does it so well.

    To do the same thing in Python (even without using the OFX module) requires a bit more effort:

    import os
    import re
    from tempfile import mkstemp
    
    filename = "test.txt"
    patternStrings = ["ZAR", r"<TD([^>]+)>([0-9]{4})([0-9]{2})([0-9]{2})"]
    replacements = ["USD", r"<TD\1>\2\4\3"]
    
    # Compile patterns
    patterns = [re.compile(p) for p in patternStrings]
    
    # Create temporary file to hold outputs
    _, tempfile = mkstemp()
    
    # Process file
    outfile = open(tempfile, 'w')
    for line in open(filename):
        for pattern, replacement in zip(patterns, replacements):
            line = pattern.sub(replacement, line)
        outfile.write(line)
    
    outfile.close()
    
    os.rename(tempfile, filename)
    

    Needless to say, the difference is even more pronounced in Java due to the large amount of boilerplate needed.

    import java.io.File;
    import java.io.FileReader;
    import java.io.FileWriter;
    import java.io.IOException;
    import java.util.Scanner;
    import java.util.regex.Pattern;
    
    public class Fixer {
        static File infile = new File("test.txt");
        static String[] patternStrings = {"ZAR", "<TD([^>]*)>([0-9]{4})([0-9]{2})([0-9]{2})"};
        static String[] replacementStrings = {"USD", "<TD$1>$2$4$3"};
    
        public static void main(String[] args) throws IOException {
            Pattern[] patterns = new Pattern[patternStrings.length];
            Scanner in = new Scanner(new FileReader(infile));
            
            // Compile patterns
            for (int i=0; i<patternStrings.length; i++) {
                patterns[i] = Pattern.compile(patternStrings[i]);
            }
            
            // Create temporary file to hold outputs
            File tempfile = File.createTempFile("tmp", "tmp");
    
            // Process file
            FileWriter outfile = new FileWriter(tempfile);
            while (in.hasNextLine()) {
                String line = in.nextLine();
                for (int i=0; i < patterns.length; i++)
                    line = patterns[i].matcher(line).replaceAll(replacementStrings[i]);
                
                outfile.write(line + System.getProperty("line.separator"));
            }        
            outfile.close();
            
            // move temp file back to filename
            tempfile.renameTo(infile);
        }
    }

    Now, at this point someone is bound to say “but wait, people don’t write desktop apps in sed!”. But that’s quite the point – Java was designed with programming in the large in mind, but it is really tedious for short programs. In many cases, learning a small domain specific language and using it to solve your small problem is faster than learning how to do the same thing in your “one language”.

    I wish I had more time to put examples together, but this situation just presented itself and got me thinking.

  • Turing completeness is a trap

    Or why you can do everything in Excel but it may not be a good idea.

    My wife’s struggle with a spreadsheet from work yesterday got me thinking about this topic again. It has perplexed me for a long time that people use Excel for so much that other tools are clearly more suited for. There are many rants on the Internet about using Excel for database-like activities, and I will probably write a bit more on that later. However, I think the key phrase in the overuse of any powerful tool is “but I can do that in my tool, too”.

    What I mean is that one could approach someone who is an expert Excel user and say “I don’t like Excel for engineering calculations because it doesn’t allow me to track the units of the numbers like for instance Frink or Mathcad does”. This guy would say “but you can do that in Excel – just do X or Y”. Somewhere they are thinking “I will use another tool when I find something that Excel can’t do”. However, they are stuck in a Turing Tarpit. The problem is that all Turing complete languages are equivalent in power in this strange and abstract sense that it is possible to do the same calculation in both. This does not say anything about how easy it will be to do that calculation, or how maintainable the code will be – these are requirements that have very little to do with the computational power.

    So, you if you are taking the view that you will use Excel as long as it is possible to do the thing you want to do, you literally don’t need any other tool. Unfortunately the same can be said about any esoteric programming language like INTERCAL or x86 assembly. It is perfectly valid to say “but I can do that calculation in INTERCAL – I don’t need any other languages”. In a strict sense this is true, but this really just points out how little Turing completeness actually buys you in terms of useful programming structures.

    At this point it seems like a good idea to mention the other thing that keeps people in Excel. I think the marketers of spreadsheets have been using the “no programming required” line for so long that everyone thinks that spreadsheets are different from programming. In fact, they can be understood as functional programming languages with a two-dimensional (or three-dimensional if you take sheets into account) structure. Of course, there’s also VBA if you aren’t sufficiently resourceful to figure out how to do everything using only the built-in functions and cells, but the spreadsheet itself is also a computational device. So a lot of people think that “using Excel” is different from “programming”, when it differs mainly in the environment.

    Bottom line: if you are a single-language kind of person (who only uses Excel or only uses C) you need to understand that it is true that you will never find a program that your favourite tool won’t be able to solve in a strict sense, but that it is also true that other tools may make it a lot easier.

  • Factors to consider when choosing a programming language

    This morning a colleague and I spoke briefly about choosing a programming language for some high-performance scientific computing (thermodynamic calculations) he wants to do. I started writing an e-mail with a couple of my thoughts, and then thought it would make a pretty good blog post for the same amount of time. So here’s my list of things to consider when choosing a programming language. Note that these are not orthogonal to one another – many seem to describe the same thing from different angles. This is just my musings, not a research study.

    1. Popularity. This is a very important one. A good place to start is the Tiobe index. You are more likely to find people to collaborate with if you use a popular language. You are also more likely to find reference material and other help. Unfortunately, the most popular language globally may not be a good match for your problem domain.
    2. Language-domain match. Choose one that matches your problem domain. You can do this by looking at what other people in your field are using (after adjusting for popularity, so don’t think the match with Java is good simply because a lot of people are using Java) or by looking at some code that solves problems you are likely to have and seeing how natural the mapping is.
    3. Availability of libraries. Some would argue that this is the same as the point above, but I don’t think so. If there’s a library that solves your problem well, you’ll put up with some ugly calling conventions or hassle in the language.
    4. Efficiency. Languages aren’t fast – compilers are efficient. Look at the efficiency of compilers or interpreters for your language. Be aware that interpreted code will run an order of magnitude slower than compiled code as a rule of thumb.
    5. Expressiveness. The number of lines of code you create per hour is not a strong function of language, so favour languages that are expressive or powerful
    6. Project-size. Do you want to be programming in the large or programming in the small? Choose a language that supports your use case.
    7. Tool support. Popularity usually buys tool support (and some languages are easier to write tools for). If you are a tool-oriented user, choose a language with good tool support. Just read this article on tool mavens vs language mavens before you make a choice.

    Note that all these things have no single right answer: they define the languages on my Pareto front. A good starting point to observe the trade-off between expressiveness and efficiency is the Computer Shootout. Also check out this analysis of some of the results.

    Of course, a personal blog need some personal input, right?

    I happen to know many languages, and the combination I use at the moment is Python for prototyping and Fortran (2003) for speed. Python is popular (it’s the second interpreted language on the Tiobe index, after PHP, which sucks for scientific stuff). It’s got a really nice set of libraries for the jobs that I am doing now and makes wrapping Fortran code easy. Fortran has some really good compilers (and the free gfortran is pretty good), and suits matrix-oriented programming really, really well. YMMV

    Both of these languages have good tool support in Emacs (my favourite text editor) and Eclipse (which I am slowly picking up).