Translated from Portuguese by Claude. Leia o original em português · more in English

note

Source Code as a Repository of Tacit Knowledge

We keep having to reflect on the impacts of artificial intelligence on software engineering practices, now that it writes code well. At one talk or another, people want a view from me, and there are two fronts that come up often and about which, to be honest, I am skeptical.

One is that generating source code will get so cheap, so simple, that you can always throw away everything that existed before. Even if that part is true, it is still hard for us to describe the theory, the definitions that exist.

Considering that we don’t know how to write down very well what the software does, that this outline is never the simplest to train, that documentation is never enough, that requirements analysis doesn’t help that much, it will be very hard to recreate the code, because a lot of the information about what the software does is contained inside it. And also in the collective unconscious of the users and of the company itself. It is no accident that an organization’s culture is so important. A bit of Conway’s law.

The other thing I am thinking about here while I wait at the restaurant is this idea that the end game is not source code. Some people say the machine will write code in machine language. It is, in theory, a performance question.

I find that very strange because it seems to me that the whole LLM mechanism, of tokens, is better suited to writing something we have abundantly documented, with plenty of examples of tests and usage. And with Python and JavaScript, we have that much more tied to the use case, to tests, to what it does and when it does it, than with a machine language, even if it is a language yet to be created.

Well, maybe a synthetic language would make more sense, but it looks like a problem we don’t need to solve, one that is already solved.

So I find it hard for us to reach this mechanism of writing a file, a project definition, a PRD, and generating the system in one shot. Simply because we don’t know how to define very well what a project is, how the software works and what it should do. It is a living organism, dependent on its users.

The other is that I also don’t believe we need to create a meta-language, actually a low-level language, that performs better, because the training mechanisms already take great advantage of these big languages, with more and more examples, even synthetic ones, of the languages themselves.

More in English

Everything in English →