I Opened a Government Open Data File With AI and Pulled Out Exactly What I Needed

AI Usage

When I heard the phrase “government open data,” I always assumed it was something difficult, and something that had nothing to do with a person like me. It is not scraped off some shady site, and it is not reserved for specialists. Today I finally opened one of those files and looked inside, together with AI.

A wide open view, standing in for opening a huge public dataset for the first time

So what is “government open data” anyway?

The file I touched this time is data officially distributed by Japan’s Ministry of Health, Labour and Welfare, the national government department that handles health and employment. A government office publishes it and says, in effect, “please, go ahead and use this,” in a form anyone can download.

It is not suspicious data. It is not data I obtained through some back channel. It is issued formally by the state, which makes it about as close to the original, trustworthy source as you can get.

Material like this is called open data, and there is a lot of it out there from public bodies. Weather, maps, statistics, all sorts of fields, published by national and local government with a note attached saying everyone is welcome to use it. What I did today was simply pick one of those files and open it.

My AI partner Kuro (Claude) put it this way.

This is data a public institution distributes officially, so you can use it with confidence. The thing is, in its raw form it is hard for a human to read. Tidying that up is exactly what I am good at.

I opened it, and it was over ten thousand rows

When I actually downloaded the file and opened it, it was quite something. More than ten thousand rows. And columns marching off sideways as well, so that at a glance I had absolutely no idea what was where.

Honestly, if I had been on my own, I think I would have taken one look and quietly closed the window. That was the reflex. Too big, forget it.

This is where AI earned its place. It swept across that enormous file and laid out for me what each column was and what kind of thing was sitting inside it. Work that would have eaten my entire day if I had followed it line by line with my own eyes turned into a rough map in no time at all.

Pulling out only the part I wanted

The moment that impressed me most was being able to cleanly extract only the part I needed out of those ten thousand-plus rows.

It turned out the data had organizing numbers assigned to it. It is the same idea as the call numbers on the spine of a library book. If you use those numbers as your handle, you can say “from this shelf to this shelf” and lift out exactly the block you want, in one clean movement.

So I asked AI to pull out only the rows under a particular number. Back came precisely the data I was aiming for, lined up neatly, nothing else.

If you search roughly by keyword, unrelated things get mixed in, or the opposite happens and you miss some. But if you slice it out by this number, you get an accurate list with nothing contaminating it and nothing left behind.

The idea came from me, not the AI

Here is the part I am quietly proud of. That approach of slicing by number was something I thought of myself, out of my own past experience.

The capability belonged to AI. The sense of the terrain, the feel for how this kind of material is organized, belonged to me. The moment those two things clicked together was the most satisfying part of the whole day.

Avoiding it because it “looked hard” was the waste

Public data was something I had avoided completely, purely on the grounds that it looked difficult. But working with AI, I got from opening the file, to understanding what was in it, to extracting the part I needed, far more easily than I had expected.

Of course, being able to pull the data out does not mean it is immediately useful for anything. What comes next is working out how to use it, and whether it is genuinely usable at all. I plan to write about that stage next time.

What I understood today is that a great many of the things we keep at arm’s length because they “look hard” are things you can walk through the entrance of, as long as you have AI with you. There is still a huge amount of solid, reliable data lying dormant out there, published free of charge by national and local government. Ordinary people are becoming able to handle it. That struck me as a quiet but genuinely large change in the age of AI.

If you want to see another day where I handed real numbers over to AI instead of guessing, there is the time I let AI read my data to find the cause of a problem. And on why I think of this as a partnership rather than following an oracle, there is AI is a partner, not a guru.

Copied title and URL