News/Views
Skip the Application, Use AI?
Agentic AI is becoming somewhat a norm. Instead of you doing something directly, you tell your agent to do it and the thing is done indirectly. This is both useful and dangerous. However, it doesn’t stop applications from still being important.
Let’s begin with that last sentence. The notion that applications (and application-producing companies, such as Adobe) will go away is dead wrong. At least in the short- to mid-term. There’s too much built up capability and performance in them for anyone to give up to a server farm somewhere to reproduce from scratch. It’s just not going to happen that you suddenly have a full AI replacement for Photoshop (or most other deep, broad applications). I’ll demonstrate that later in the article.
To understand AI today you need to also understand the dicotomy between centralized and decentralized computation and data. The big tech companies will always push centralization, mostly because that gives them control over you. They have your data. Only they can operate on it. They control what you can and can’t do with it. They charge you for dealing with your own information ;~). We consumers, however, prefer decentralization. We want to own and control our own data and do with it what we will. We don’t want to keep paying someone else for things we can do ourselves. Mainframes were centralization. Personal computers were decentralization.
When things are centralized, you need more communications bandwidth. When things are decentralized you need more computational power in the individual devices. We’ve been in a constant push/pull with the whole centralization issue from the very beginning of computing, but note one thing about today: billions of people (and small-to-big companies) have their own devices and (mostly) control them. Trying to tell these folk they can only use a giant mothership of a server to do anything is Orwellian, and ultimately not likely to work. You may believe that the future is dystopian due to all the dystopian media these days, but note that the central theme of most dystopian movies and television shows is that individuals rise up to change the situation. That would be you, dear reader (except for you Elon ;).
With that out of the way, let’s talk about my thesis here. Basically that’s AI will not get rid of applications any time soon.
I have a pretty long history with AI dating back to the 1980’s, so I’ve watched how it's progressed. Currently we’re in a Silicon Valley-style Race-to-Dominate cycle with AI. Profits be damned, it’s who gets the most customers first that will ultimately win. We’ve seen how this plays out before, particularly with the so-called Internet 2.0 wave that crashed in 2000. There’s a handful of well-funded players trying to lock up users with the free-sample-now-pay-me-later approach. If you haven’t read Cory Doctorow’s Enshitification, do so now, as it will tell you what happens next.
Meanwhile, we have this media-hyped idea that AI will do everything, and you won’t need apps any more. What mainstream media hasn’t recognized yet is that those of us that employ AI (see my personal use statement, below) don’t use it that way. Most of us use AI within applications or externally to instruct applications. This is the infamous “Can you magnify that image and pull out more details?” approach that television procedurals have made famous. No one is intensely writing code and using sophisticated UI in those programs, they just press a button and something somewhere in the machine they’re looking at does the job (e.g. AI).
I’ve been using that type of approach with Xcode (programming), Elements (Web site design), and Excel (spreadsheets) for some time now. I know what I want to do, and I know it will often take a lot of typing to do it. Indeed, that’s generally the crossover point: saying what I want done is simpler than doing it. Thus, telling an AI agent to “give this spreadsheet a heading, style it with my current style, apply my logo, check all the formulas and spelling, and make sure the column sizes are appropriate and matched wherever possible” takes me all of a few seconds to say (or type). Minutes later, my AI will have done exactly that, I’ll review that it did what I wanted, and I end up saving a fair bit of time. This is not AI doing everything (as in “from scratch, create a final document all by yourself that has columns and rows with this data and performs these operations on the data,” etc.).
What’s happening in this AI-handoff approach is distributed processing. My local computer is running Excel, which recognizes commands and executes them, while some server somewhere (currently moved to my desktop on another computer, and sometimes living on the same laptop as I’m working on* ;~), tells Excel to do a series of things based upon its knowledge of how Excel works and what spreadsheets should look like.
Now, you might think that’s all well and good, but what does that have to do with photography? Simple: Adobe has recently opened up their apps to agentic AI. So you could say or type “adjust the image currently opened in Photoshop so that the shadows are lifted, skin tones are kept intact, but saturation is increased some.” That’s a chain of things you’d normally do with multiple different UI controls within Photoshop. The chain and complexity can get quite large with some requests, which means that if you had tried to do this in Photoshop itself you’d have needed to know where all the appropriate controls live, how they work, how to adjust them and control for interdepencies, and so on. Generally this style of work becomes useful when it takes less time to tell the AI what to do and review what it did compared to you doing it yourself. It’s also very valuable for batch work. As in a wedding photographer instructing an AI like this: “Image A has had all my adjustments I want applied. Make all the other images in folder X match this look. Pay careful attention to keeping skin tones, colors, and background blurs consistent, apply only the level of skin smoothing and sharpening you find in Image A, and if you believe an image should be cropped, confirm the crop with me.” That would be a lot of work to do image by image by hand. But it’s easily done by AI these days, though you have to keep a careful watch over results. Think of AI as an assistant. Manage it as you would an assistant.
You might be seeing why I say that applications won’t go away any time soon. They’re a useful tool. Not just for individual users, but for AI, too. Apple seems to understand this in the way they’re approaching AI in their OS’s. It’s more agentic and specific AI agnostic, processed on device much of the time, and allows deep application integration (if the application developer integrates it).
The more sophisticated the application—Photoshop, Final Cut Pro, Excel, et.al.—the more it’s likely to stick around. Application developers become the repository for creating sophisticated specific tool actions (and interactions), AI becomes a way of instructing the tools in the applications. But even simple applications aren’t likely to go away, because by definition a simple application does one or a few thing(s) well and directly, and it’s faster for a user to just do that directly. It’s only the applications that live in the gray area in between simple and sophisticated that I’d worry about going away, as they’re not simple enough to be directly user controlled and not sophisticated enough to expose useful tools to AI.
You’ll note that I haven’t mentioned generative AI. That’s a different thing. AI is pattern-based. It was trained on patterns, it outputs patterns. The early AI application in photography was generative and generally fell into one of three categories: (1) take out recorded anomalies (e.g. noise) that doesn’t belong; (2) correct elements that are compromised (e.g. sharpen edges that have been aliased); and (3) replace or add a section of the image with pixels that reflect the rest of the image. The first two are very useful, and better than the brute force tools we used to use for noise reduction and sharpening. The third is controversial because the resulting pixels were not generated photographically. That said, on occasion I’ll use #3 on a background distraction that I couldn’t control in the field and then evaluate whether the result is likely what I would have gotten had I been able to remove the background distraction in the field. If yes, then I’ll allow it. If no, I’ll just live with the distraction. And that’s actually no different than I did in the darkroom (yes, you can add/remove things on the enlarger, but that was a hellish task, so mostly avoided).
Note that I’m writing about AI fairly narrowly here. Specifically, as it currently and is likely to apply to photography. The whole “AI will destroy the world” doomsaying going on in mainstream media is a whole different story, and not photographic specific. I’m talking more Topaz here than ChatGPT. In other words, targeted tools versus generalized do-anything ones. Targeted AI tools are useful and aren’t going to make our best applications go away.
———
* You probably want to know how I do that. I use LMStudio and an open-AI model that is appropriate for the job at hand. Both are free, but both require a large amount of local device horsepower (Mac users think M3 with 32GB RAM or better, preferably current processor and 64GB or more). As for model, dozens are available. Some are targeted at coding, some at math, some at words. I’ve got several loaded in my machines. To make a local AI work with a language, both need to support MCP (Model Control Protocol), which will probably take some geeky setup to get right, but once connected, I can type “In Project byThom MAX add a page for Flexible Picture Controls (Recipes). Populate it with my header and footer, set a Container in the middle that includes cards that I’ll populate with individual data later, so use a placeholder image and text.” When I do that in my AI, it uses the MCP connection to step Elements through all the things implied by what I typed. A minute or two later I evaluate what it did and start making small changes to the results.
———
Disclosure: Does Thom (and byThom) use AI? The answer to that is a bit complicated, but let’s start with an easy point: if you’re reading something on byThom or by me, it was typed by my fingers from instructions from my brain. If you see a photograph on byThom or by me, it was taken with a camera on instructions/settings from my brain, and processed with instructions from my brain. If I were ever to use an AI-generated statement or photo on my sites, it would labeled as one.
However, I do use AI for some work. For example, I use it to instruct XCode (for program development), Elements (for Web site design), and Excel (for spreadsheet verification and reformatting). Generally I’m doing the same things with AI I would do with a human staff: delegating a specific job to be done to my goals and specifications, and reviewing it once it is done. So when I designed my first new app (coming soon, with more following), I did a full MRD (Marketing Requirements Document), a full mockup (design shell), specified an architecture and coding approach, wrote a full manual (!), and then and only then handed off some function-creation to Claude. Upon Claude’s completion of a task, I review it. Sometimes I adjust something, sometimes I scold Claude, sometimes I accept what it did, all pretty much what I used to do with a human staff of programmers. As most programmers using AI now do, I have Markdown documents that contain the plan, concerns, testing requirements, and much more, and both Claude and I are constantly conversing in and adjusting those documents. Did I write a lot of the code for the first app I’m going to launch? No, though I did some, particularly in the UI area. Did I doublecheck everything the AI created? Absolutely.
Finally: yes, I will use AI-controlled noise reduction and sharpening on my images at times. I’m convinced that AI does a better job than the previous brute-force tools we had for those things, without introducing new problems or incorrect/misleading information. But like any tool, I use these in moderation, not extremes.
I Was Wrong
Back in 2011 I told Nikon executives that it was not useful to pursue a relationship with a phone maker. Others at the time, particularly Leica, were doing so. My contention was that the overall sales and margins on the R&D and supply for the small sensor/lens combo would be too small to pursue. Just licensing the Nikon name would have been a better economic choice, and there you’re still only talking about peanuts.
Meanwhile, I believed that the camera-as-lens attachment ideas both Olympus and Sony pursued wouldn’t work either, because they had a dual problem: (1) they had to control said camera-as-lens from an app; and (2) communication would be via bluetooth/Wi-Fi, which is notoriously unreliable as a full-time solution. #1 meant that the result would be an add-on and live on top of the software stacks, undermining performance. #2 meant fiddly settings and lagging connectivity.
I’ll still stand by those statements, but I wasn’t thinking enough out of the box (unusual for me). Now that we’re in a world of magnetic connections and USB-C, the work-with-smartphone maker problem becomes far simpler, and opens up opportunities that weren’t there in 2011. Simply put, add a USB-C port to the MagSafe ring. This gives you both physical and electronic attachment. MagSafe rings are big and strong enough to support an m4/3 or APS-C camera-as-lens, for sure, producing an upsell for the phone user who wants to offer better imaging to their top customers.
The camera company would still need to have good low-level engineering access within the phone company, thus it would have to be a partnership type agreement, but this could be a win-win for both parties. My problem in 2011 was that I couldn’t see any win-wins, only win-loss, loss-loss, or loss-win.
The only problem with the USB MagSafe approach is that it may be too late to have meaningful traction today. Image sensors in phones have gotten bigger, and they’re now firing off with multiple focal lengths, too. Gimbal cameras sit right above what the phones do, and dedicated compact cameras are making a comeback. Still, it’s worth contemplating how you attract and hold the young. They start with smartphones, so how do you get them across into the dedicated camera world eventually? We’ve seen the camera companies try that with vlogging type cameras with modest success, particularly Sony. But when everyone is doing that, how do you stand out? Well, you could stand out by both having a long legacy in cameras as well a significant smartphone connection. That pretty much leaves Sony as the only company that could pursue the MagSafe camera add-on by itself, but I think the failure of the QX line probably has them loathe to try that camera-as-lens approach.
I’m only writing this to be consistent. I’ve written about how I was right about some ideas before anyone else thought of them. So I need to also write about when I’ve been wrong. No one’s perfect at understanding how to reach the best future. Sometimes the path has a few twists and turns, as well as dead-ends and cul-de-sacs.