Year: 2010

  • Impressions of superiority

    None of my compliments are as often or firmly rebuked as telling someone they have a natural ability for something.

    It is a quirk of our society that our sense of achievement is diminished by aptitude. We praise above all individuals who have overcome adversity in their achievement. Many view being told they have an aptitude or ability as tantamount to saying they didn’t have to overcome as much as other people would have.
    I am aware of my natural aptitudes for some activities. I had no particular difficulty learning to play the guitar and appear to understand music pretty well. My mathematical intuition is sound. I find it easy to learn new programming languages. I also have certain… challenges. I am easily distracted, I find medium-term planning difficult and my learning mode is often highly inefficient as I require a lot of background and overview before things start making sense. It has taken me a lot of time to figure out what the population average abilties in some of these areas are. I am by no means there, but I have learned enough to know which properties are high and which are low.
    The reason for listing my own abilities here is an attempt to defuse the knee-jerk reaction that I have perceived many times when you tell other people they have abilities above the norm. Many people will vehemently deny such abilities because they assume you are implying they didn’t have to work at it. It seems more acceptable to lament the lack of such abilities in others than to claim them for yourself. Many people don’t even do that – they will deny any particular variation in these abilities among people and instead end up lamenting their lack of application. I find this particularly objectionable, as the implication is often that the reason for failure is fundamentally a weak will.
    Here’s how a sample dialogue often plays out:
    Me: I have left off marking my papers again, and will have to put in a late night tonight.
    Them: Oh, I always get my marking out of the way first thing, then I feel much better.
    Me: I wish I had your well-developed work ethic and drive to completion
    Them: I don’t really have any particularly well-developed drive to completion – you can do this too.
    Me: I have tried and failed. Perhaps I am just weak-willed. But more likely I am more easily distracted than you are, or find this task more onerous than you do
    Them: How dare you imply that I don’t need to apply myself to get this stuff done!
    Unequal distribution of abilities sits poorly with the doctrine of equality (or at least equal opportunity). Even worse, it is very difficult to think of oneself as anything other than “normal”. When I lost 25 kg, I honestly did not seem any different to myself in the mirror. I can just remember seeing a lot more fat people around. When I started running, it never felt much different to run at a comfortable pace, even though that has changed radically as I trained. It’s an all-too easy trap to fall into as ones abilities develop: renormalisation. As you become more knowledgable, you just start to see more idiots. As you become less fit, you start wondering how all these people can be so active.
    This post is a bit of an impression rather than a well-though out argument. I wanted to get this down while I was still thinking about it after having discussed it at a party last night. There is much left out here: how much of these abilities are innate and how much are learned? How does practice feature into it and how well can we assess our own abilities, and what is the mechanism of self-control (in fact, what is being controlled). That will have to wait for later.
  • Lecturer rant reprise

    My post yesterday prompted many comments. This post attempts to address some of them.

    I got a couple of people expressing the idea that my experiences with suboptimal systems are not unique to my institution. I appreciate your comments, especially those mentioning specific institutions, as they give a bit of insight into the scale of the problem. I find widespread problems that would be (relatively) simple to solve in a (relatively) general way incredibly irksome. I feel great desire to solve these problems and distribute the solutions to everyone. The way I feel is quite well described by the hacker ethic. I aim to do meaningful work, but many of the things I am forced to do at the university are menial work. It gives me little consolation that others are doing the same things. In fact, it makes it worse.
    My boss chided me for putting the post online for everyone to see. His concern was for airing dirty laundry in public. I took the university’s name out of the post but I stand by what was said and I don’t think much is going to happen due to that rant one way or another, save perhaps serving as an aide memoire for the marking system ideas and a nucleation point for further discussion about it. The whole wikileaks furore was also being discussed copiously online at the time and I was struck by the urge for secrecy. I understand the need for privacy and feel the urge to confide in particular people myself, but I don’t feel quite the same way about organisations. I tend to feel they have a responsibility to be transparent. What is wrong with the idea that by talking openly about problems, you increase the likelihood of them being solved? Especially if many places are having the same problems?
    Later that day I had a bit of back-and-forth on Twitter with a mate who worked in the lab with me several years ago. He asked what I had done about the problems and suggested that I develop a prototype of a system to show to admin. His posts hit a nerve, because this is exactly the attitude that I have had up to now (don’t complain – just fix it). In the years I have been at the university I have developed the following tools that led out from me suggesting that things could be done better (I list my estimated cumulative time investment into these systems and the number of years they have been in development in brackets):
    • A static web page generator that we used to generate our departmental webpage for a couple of years before the new CMS was rolled out (100 hours, 4 years)
    • A system that uses the above along with a Google Docs spreadsheet to collect information about the department’s publications and generates a cross-linked static website suitable for burning to CD (50 hours, 3 years)
    • A project allocation system that uses GLPK to assign projects to students for our fourth-year project activity based on their preferences (150 hours, 3 years)
    • Scripts for typesetting the timetables for the school of engineering based on a flat format spreadsheet generated from syllabus. (250 hours, 4 years)
    • VBA code for use in Excel to calculate student number checksums and several shared calculations, including the supp calculation I mentioned in the previous post (10 hours, 2 years).
    • Roll-out of a departmental Wiki to capture information about these and other projects along with other documentation (50 hours 3 years)
    This list is not meant to sing my praises, but I suppose it serves to establish that I do not simply moan about things – I generally just say “that can be done better” and then move to do things better. The problem with this attitude is that these things are way outside even the most liberal reading of my job description, which means that all the time I spent on these projects is time that I could have (should have?) spent on my students, or developing better lectures or my PhD or any of the things that are actually expected of me.
    I completely understand why no-one is doing these things: because it is no-one’s job. Now, I get a bit annoyed at the “fix it yourself” attitude because I encounter it most often when it has anything to do with computers. I am a relatively capable programmer, and often when I complain that I can’t find a particular tool or about how a particular tool works I get this glib answer that complaining doesn’t solve anything and I should do something about it myself or shut up.
    Imagine a scenario: I post a tweet complaining that water is leaking from the ceiling of my office onto my keyboard. I decry the poor maintenance policies that led to the leak occurring and point out that a system should be in place which allows me to report such incidents and have them resolved in a timeous fashion. Would it be reasonable to ask me to develop a maintenance schedule for campus and perhaps implement a call-center? I think not, because such activities clearly belong with facilities management.
    I can think of endless cases where I lack the skills or the time to fix things that bother me. I have already put in a lot of time developing half-baked solutions to problems that could be solved far better by people who are actually trained in this work. I am not a programmer, I am not a systems analyst.
    Let’s say I want to build a house. I think it is reasonable to approach an architect with a functional spec (I want a nice kitchen, I want the house to be cool, I like this kind of design) and expect him to develop a detailed design rather than drawing up plans for the house myself. He is bound to do a much better job because this is what he is trained to do. I think it is similarly reasonable to point out the gaps in the admin system and expect them to either streamline the system themselves or at least bring in someone who can.
  • Lecturer rant: there should be a system for marks

    Even after six years at the university I am often amazed at how stupid we are.

    I use “we” there because I am part of faculty. I am amazed because of these stupid things (and these are just the ones I know of):

    Yearbooks

    The university yearbooks are written and maintained by hand. Yep, if you clicked on some of those links and saw lists of repetitive seemingly machine-generated content, you may infer that this is all on a database somewhere and simply sucked in to this document by some reporting system. You would be wrong. There are people in each department tasked with maintaining the system. Changes to the document are posed in the form of a manually generated diff and applied by humans. This means there is no automated consistency checking and no way of generating the dependency tree of a particular subject to answer a question like “I have failed GFU 320 – does this cost me a year?” because the dependencies are simply listed in the document. I have written scripts that scrape the PDFs for useful information like the number of periods allocated etc to check that they are consistent with the timetable (see the next point) and generate graphs like the one below that show the dependencies of our course. Why is there not an automatically generated graph like this with clickable links into the subject descriptions on the UP website? The graph is generated using Graphviz

    Timetables

    There is no single, up-to-date timetable for the university as a whole. Once a year the timetable book gets generated from a program called Syllabus Plus. After this point, all the advanced features of Syllabus, including distributed access by multiple people, custom web-based timetable views etc are ignored and the Word file containing the manually reformatted timetable information sorted by subject code is sent to each department. Each member of faculty then has to traverse this document searching for their subject and draw up a timetable in the traditional two-dimensional grid form themselves. This is a huge waste of effort! I have written scripts that read a dump prepared from Syllabus by our helpful IT staff in an Excel file that changes subtly each year even though I have given a detailed spec of the format I require and prepares timetables for the school of engineering.

    Unfortunately, this is not the end, because the venue bookings are handled on a totally different system which does not write back to Syllabus. This means that at any given moment it is literally impossible to reconstruct the exact timetable as it stands without talking to every lecturer involved, as the ad-hoc changes are only relayed to them and not recorded centrally beyond the subject name (not the group or other information you may need to generate the timetable). The whole situation is silly and costs everyone involved hours of their time as they sift through the book trying to verify that the information is correct. What makes it worse is that several other universities are allowing staff access to the syllabus plus system while we are using what boils down to one step above manual scheduling.

    Student marks

    Student marks are calculated by each lecturer using home-rolled spreadsheets that are almost guaranteed to have an error in somewhere. The rules to determine for instance if a student qualifies for a supp are subtle and often incorrectly implimented. To give you an idea, here is the logic for if a student qualifies for a supp in the school of engineering from the 2010 yearbook:

    In the School of Engineering a supplementary examination is only granted in instances where:

    1. A final mark of between 45% and 49% was achieved;
    2. A final mark of between 40% and 44% was achieved and where the candidate also achieved either a semester mark or an examination mark of 50% or higher;
    3. A pass mark has been obtained, but the required subminimum in the examination section of the module or divisions thereof has not been obtained.
    4. A final mark of between 40% and 49% has been obtained in first-year modules on 100 level.
    I have made a graph showing the regions for a 50% semester mark exam split:

    Is it not glaringly obvious that this should at least be coded up in a reusable spreadsheet function that can be distributed to everyone? Of course, that’s just a sticking-plaster, because the system is broken in more subtle ways.
    At the beginning of each course, an automatically-generated class list is sent to you by admin. It used to be a CSV file, now it is an Excel file (already my scripting has taken a knock right there). You then calculate marks in any way you want and send the marks back. Unfortunately they now send you a revised class list which does not have the same students in the same order as the previous ones. So what could have been a simple copy-and-paste now either becomes tedious lookup calculations or even more tedious manual lining up of the data. I think it is entirely obvious that we need things like the following ideas:
    • Accelerated marking system you scan in the memo and the papers, define active regions on the memo and mark directly into the system. The papers have a barcode sticker from a pile you issue the students at the beginning of each year, so your (lecturer) workflow involves going into the system and clicking through the memo for each paper. Totally transparent to the student and externals, a lot faster than carting around tons of papers and shuffling through them.
    • Grade calculation based on a tree structure with calculation nodes most all grade calculation can be seen as a tree-like aggregation of primary grades. Final mark is weighted average of semester and exam, each paper is sum of questions, which is sum of subquestions etc. You need a builder and navigator, with colour highlighting and live stats to answer questions like ‘why is this guy’s semester mark so low’ or ‘what will my class average look like if I adjust my semester mark’.
    • Audit trails all changes should be historied (BIG fail with spreadsheet-based systems). If an adjustment is made, this should be noted for this student so that one can track the total ‘free marks’ a particular student has received. This can also be used in
    • Automatic edge case detection: Students that are ‘close’ to particular boundaries like distinctions or failing or supps should be identified by the system and automatic adjustments should be suggested. This would of course be historied so that you can review how much action has been taken. With a proper system, this would include all the student’s subjects, so that the system would warn you if a student (for instance) only needs your subject to pass his whole course and has not received a lot of free mark. I have even thought a good system could be that you start your course with 0.5 percent per test (so 0.5 for semester mark, 0.5 for exam and 0.5 for final mark), so 1.5 per course. This is the adjustment we allow today anyway, as we round each of those marks. Lecturers can use some of these marks for adjustments. The system could even find optimal distribution of the free marks to maximize the chance that the student succeeds in the course. I mean, we’re only probably repeatable within about 2% to 5% on marks anyway.
    • Aggregated statistics lecturers should be able to see the distributions of the other courses this group has, and be warned about strange anomalies like an otherwise good student scoring really low or an average student doing particularly well.
    The thing that amazes me most is that we have all this student manpower doing projects to design websites and database systems and intelligent algorithms for fault-finding and inference, but we’re not using any of it on our own systems! I am sure that the development of a system with some of the features above would not be beyond our student body. A part of the problem is that no-one sees development of such systems as part of their job. If one person were to try to develop such a system they would end up wasting a lot of their own time, as the total loss of time is distributed among so many people. This means there should be a higher-up decision about introducing more efficiency in the system rather than a lone developer working on it. Unfortunately, the small losses aren’t accounted for in an actionable way and things keep going like they’re going. Perhaps I will should send a link to this blog post to my head of department…
  • Turing completeness is a trap

    Or why you can do everything in Excel but it may not be a good idea.

    My wife’s struggle with a spreadsheet from work yesterday got me thinking about this topic again. It has perplexed me for a long time that people use Excel for so much that other tools are clearly more suited for. There are many rants on the Internet about using Excel for database-like activities, and I will probably write a bit more on that later. However, I think the key phrase in the overuse of any powerful tool is “but I can do that in my tool, too”.

    What I mean is that one could approach someone who is an expert Excel user and say “I don’t like Excel for engineering calculations because it doesn’t allow me to track the units of the numbers like for instance Frink or Mathcad does”. This guy would say “but you can do that in Excel – just do X or Y”. Somewhere they are thinking “I will use another tool when I find something that Excel can’t do”. However, they are stuck in a Turing Tarpit. The problem is that all Turing complete languages are equivalent in power in this strange and abstract sense that it is possible to do the same calculation in both. This does not say anything about how easy it will be to do that calculation, or how maintainable the code will be – these are requirements that have very little to do with the computational power.

    So, you if you are taking the view that you will use Excel as long as it is possible to do the thing you want to do, you literally don’t need any other tool. Unfortunately the same can be said about any esoteric programming language like INTERCAL or x86 assembly. It is perfectly valid to say “but I can do that calculation in INTERCAL – I don’t need any other languages”. In a strict sense this is true, but this really just points out how little Turing completeness actually buys you in terms of useful programming structures.

    At this point it seems like a good idea to mention the other thing that keeps people in Excel. I think the marketers of spreadsheets have been using the “no programming required” line for so long that everyone thinks that spreadsheets are different from programming. In fact, they can be understood as functional programming languages with a two-dimensional (or three-dimensional if you take sheets into account) structure. Of course, there’s also VBA if you aren’t sufficiently resourceful to figure out how to do everything using only the built-in functions and cells, but the spreadsheet itself is also a computational device. So a lot of people think that “using Excel” is different from “programming”, when it differs mainly in the environment.

    Bottom line: if you are a single-language kind of person (who only uses Excel or only uses C) you need to understand that it is true that you will never find a program that your favourite tool won’t be able to solve in a strict sense, but that it is also true that other tools may make it a lot easier.

  • Life as a data source

    Tracking progress is a powerful mechanism for improvement, as we found out with our weight loss attempts the year before we started running. Using the power of feedback (I am a control man after all), I was able to reduce my weight and keep it off. My daily weight measurements are a part of my routine and I just love analysing the stats.
    This graph shows my weight since I’ve started weighing myself at the end of November 2008. I estimate that I weighed above 100 kg when we started with the weight loss, but the first measurements are around 95 kg. This means I lost between 20% and 25% . Most of that loss was linear, around 1 kg/week as predicted by a 4600 KJ/day dietary energy deficit. As a rule of thumb, fat has about 32MJ/kg stored energy, so to lose a kg of fat you need to have that much energy deficit. These figures are approximately converted from the Hackers diet, which is a really nice reference for weight loss.
    We have nice scale that also measures body fat via an impedance measurement. I don’t trust the fat percentage much as an absolute thing, as it is highly correlated to the measured weight, as shown here:
    However, it does seem to track my more subjective gauge of flabbiness, and as any control person will tell you, measurements that get the direction right are already helpful, even if they aren’t accurate in an absolute sense. This is especially true here, as the target for body fat is just as low as possible (since it is highly unlikely that I will be able to get it down to anything approaching unhealthy levels).
    Many people advocate exercise for weight loss. When you consider the kind of energy deficits you need to really lose weight, my personal take is that it’s easier to eat less than to work out more. In fact, when you’re losing, your energy levels go down so you’re even less likely to work out. What worked for me was to lose most of my weight and start exercising when I was lighter.
    Based on the success of tracking my weight, when I started running, I initiated the activity by buying a Garmin ForeRunner 305 so that I could track my progress. It turns out that the Mac software stores all the data in a sqlite database, so I whipped up some some python code that pulls the stuff out of there and analyses it. It’s hosted on Github. A lot of the time spent writing the code was put into writing a routine that removes all the duplicate entries from the database. I’ve tried getting Garmin to do something about it by posting a request on their forums, to no avail.
    After cleaning up the database, the scripts take each run and approximates the speed from the displacement data. Everything is resampled to one second intervals to make data processing easier. I started out plotting everything with gnuplot – at some point I suppose I should make the move to matplotlib for all the plotting, but separating the plotting process is a bit easier with gnuplot and a Makefile.
    The results: Here’s a speed histogram, showing the distribution of speed for each run. This is handy to see if for instance training for higher speeds is working.
    The representation I am most proud of is this graph:
    This shows the pareto frontier of my running performance over speed and distance. On the top right the world records are shown, and I also plot a percentage of the world record speed as a target. That percentage is at the moment calculated from a 20 minute 5 km target.
    If people had told me how much data could be extracted from running, I may have started a lot sooner. Life is a great data source