e-Literate

Present is Prologue

Tag: AI/ML

  • ChatGPT: Post-ASU+GSV Reflections on Generative AI

    ChatGPT: Post-ASU+GSV Reflections on Generative AI

    The one question I heard over and over again in hallway conversations at ASU+GSV was “Do you think there will be a single presentation that doesn’t mention ChatGPT, Large Langauge Models (LLMs), and generative AI?”

    Nobody I met said “yes.” AI seemed to be the only thing anybody talked about.

    And yet the discourse sounded a little bit like GPT-2 trying to explain the uses, strengths, and limitations of GPT-5. It was filled with a lot of empty words, peppered in equal parts with occasional startling insights and ghastly hallucinations. 

    That lack of clarity is not a reflection of the conference or its attendees. Rather, it underscores the magnitude of the change that is only beginning. Generative AI is at least as revolutionary as the graphical user interface, the personal computer, the touch screen, or even the internet. Of course we don’t understand the ramifications yet.

    Still, lessons learned from GPT-2 enabled the creation of GPT-3 and so on. So today, I reflect on some of the lessons I am learning so far regarding generative AI, particularly in EdTech.

    Generative AI will destroy so we can create

    Most conversations on the topic of generative AI have the words “ChatGPT” and “obsolete” in the same sentence. “ChatGPT will make writing obsolete.” “ChatGPT will make programmers obsolete.” “ChatGPT will make education obsolete.” “ChatGPT will make thinking and humans obsolete.” While some of these predictions will be wrong, the common theme behind them is right. Generative AI is a commoditizing force. It is a tsunami of creative destruction.

    Consider the textbook industry. As long-time e-Literate readers know, I’ve been thinking a lot about how its story will end. Because of its unusual economic moats, it is one of the last media product categories to be decimated or disrupted by the internet. But those moats have been drained one by one. Its army of sales reps physically knocking on campus doors? Gone. The value of those expensive print production and distribution capabilities? Gone. Brand reputation? Long gone. 

    Just a few days ago, Cengage announced a $500 million cash infusion from its private equity owner:

    “This investment is a strong affirmation of our performance and strategy by an investor who has deep knowledge of our industry and a track record of value creation,” said Michael E. Hansen, CEO, Cengage Group. “By replacing debt with equity capital from Apollo Funds, we are meaningfully reducing outstanding debt giving us optionality to invest in our portfolio of growing businesses.”Cengage Group Announces $500 Million Investment From Apollo Funds (prnewswire.com)

    That’s PR-speak for “our private equity owners decided it would be better to give us yet another cash infusion than to let us go through yet another bankruptcy.”

    What will happen to this tottering industry when professors, perhaps with the help of on-campus learning designers, can use an LLM to spit out their own textbooks tuned to the way they teach? What will happen when the big online universities decide they want to produce their own content that’s aligned with their competencies and is tied to assessments that they can track and tune themselves? 

    Don’t be fooled by the LLM hallucination fear. The technology doesn’t need to (and shouldn’t) produce a perfect, finished draft with zero human supervision. It just needs to lower the work required from expert humans enough that producing a finished, student-safe curricular product will be worth the effort. 

    How hard would it be for LLM-powered individual authors to replace the textbook industry? A recent contest challenged AI researchers to develop systems that match human judgment in scoring free text short-answer questions. “The winners were identified based on the accuracy of automated scores compared to human agreement and lack of bias observed in their predictions.” Six entrants met the challenge. All six were built on LLMs. 

    This is a harder test than generating anything in a typical textbook or courseware product today. 

    The textbook industry has received ongoing investment from private equity because of its slow rate of decay. Publishers threw off enough cash that the slum lords who owned them could milk their thirty-year-old platforms, twenty-year-old textbook franchises, and $75 PDFs for cash. As the Cengage announcement shows, that model is already starting to break down. 

    How long will it take before generative AI causes what’s left of this industry to visibly and rapidly disintegrate? I predict 24 months at most. 

    EdTech, like many industries, is filled with old product categories and business models that are like blighted city blocks of condemned buildings. They need to be torn down before something better can be built in their place. We will get a better sense of the new models that will rise as we see old models fall. Generative AI is a wrecking ball.

    “Chat” is conversation

    I pay $20/month for a subscription to ChatGPT Plus. I don’t just play with it. I use it as a tool every day. And I don’t treat it like a magic information answer machine. If you want a better version of a search engine, use Microsoft Bing Chat. To get real value out of ChatGPT, you have to treat it less like an all-knowing Oracle and more like a colleague. It knows some things that you don’t and vice versa. It’s smart but can be wrong. If you disagree with it or don’t understand its reasoning, you can challenge it or ask follow-up questions. Within limits, it is capable of “rethinking” its answer. And it can participate in a sustained conversation that leads somewhere. 

    For example, I wanted to learn how to tune an LLM so that it can generate high-quality rubrics by training it on a set of human-created rubrics. The first piece I needed to learn is how LLMs are tuned. What kind of magic computer programming incantations do I need to get somebody to write for me?

    As it turns out, the answer is none, at least generally speaking. LLMs are tuned using plain English. You give it multiple pairs of input that a user might type into the text box and desired output from the machine. For example, suppose you want to tune the LLM to provide cooking recipes. Your tuning “program” might look something like this:

    • Input: How do I make scrambled eggs?
    • Output: [Recipe]

    Obviously, the recipe output example you give would have a number of structured components, like an ingredient list and steps for cooking. Given enough examples, the LLM begins to identify patterns. You teach it how to respond to a type of question or a request by showing it examples of good answers. 

    I know this because ChatGPT explained it to me. It also explained that the GPT-4 model can’t be tuned this way yet but other LLMs, including earlier versions of GPT, can. With a little more conversation, I was able to learn how LLMs are tuned, which ones are tunable, and that I might even have the “programming” skills necessary to tune one of these beasts myself. 

    It’s a thrilling discovery for me. For each rubric, I can write the input. I can describe the kind of evaluation I want, including the important details I want it to address. I, Michael Feldstein, am capable of writing half the “program” needed to tune the algorithm for one of the most advanced AI programs on the planet. 

    But the output I want, a rubric, is usually expressed as a table. LLMs speak English. They can create tables but have to express their meaning in English and then translate that meaning into table format. Much like I do. This is a funny sort of conundrum. Normally, I can express what I want in English but don’t know how to get it into another format. This time I have to figure out how to express what the table means in English sentences.

    I have a conversation with ChatGPT about how to do this. First I ask it about what the finished product would look like. It explains how to express a table in plain English, using a rubric as an example. 

    OK! That makes sense. Once it gives me the example, I get it. Since I am a human and understand my goal while ChatGPT is just a language model—as it likes to remind me—I can see ways to fine-tune what it’s given me. But it taught me the basic concept.

    Now how do I convert many rubric tables? I don’t want to manually write all those sentences to describe the table columns, rows, and cells. I happen to know that, if I can get the table in a spreadsheet (as opposed to a word-processing document), I can export it as a CSV. Maybe that would help. I ask ChatGPT, “Could a computer program create those sentences from a CSV export?” 

    “Why yes! As long as the table has headings for each column, a program could generate these sentences from a CSV.” 

    “Could you write a program for me that does this?” 

    “Why, yes! If you give me the headings, I can write a Python program for you.” 

    It warns me that a human computer programmer should check its work. It always says that. 

    In this particular case, the program is simple enough that I’m not sure I would need that help. It also tells me, when I ask, that it can write a program that would import my examples into the GPT-3 model in bulk. And it again warns me that a human programmer should check its work. 

    ChatGPT taught me how I can tune an LLM to generate rubrics. By myself. Later, we discussed how to test and further improve the model, depending on how many rubrics I have as examples. How good would its results be? I don’t know yet. But I want to find out. 

    Don’t you?

    LLMs won’t replace the need for all knowledge and skills

    Notice that I needed both knowledge and skills in order to get what I needed from ChatGPT. I needed to understand rubrics, what a good one looks like, and how to describe the purpose of one. I needed to think through the problem of the table format far enough that I could ask the right questions. And I had to clarify several aspects of the goal and the needs throughout the conversation in order to get the answers I wanted. ChatGPT’s usefulness is shaped and limited by my capabilities and limitations as its operator. 

    This dynamic became more apparent when I explored with ChatGPT how to generate a courseware module. While this task may sound straightforward, it has several kinds of complexity to it. First, well-designed courseware modules have many interrelated parts from a learning design perspective. Learning objectives are related to assessments and specific content. Within even as simple an assessment as a multiple-choice question (MCQ), there are many interrelated parts. There’s the “stem,” or the question. There are “distractors,” which are wrong answers. Each answer may have feedback that is written in a certain way to support a pedagogical purpose. Each question may also have several successive hints, each of which is written in a particular way to support a particular pedagogical purpose. Getting these relationships—these semantic relationships—right will result in more effective teaching content. It will also contain structure that supports better learning analytics. 

    Importantly, many of these pedagogical concepts will be useful for generating a variety of different learning experiences. The relationships I’m trying to teach the LLM happen to come from courseware. But many of these learning design elements are necessary to design simulations and other types of learning experiences too. I’m not just teaching the LLM about courseware. I’m teaching it about teaching. 

    Anyway, feeding whole modules into an LLM as output examples wouldn’t guarantee that the software would catch all of these subtleties and relationships. ChatGPT didn’t know about some of the complexities involved in the task I want to accomplish. I had to explain them to it. Once it “understood,” we were able to have a conversation about the problem. Together, we came up with three different ways to slice and dice content examples into input-output pairs. In order to train the system to catch as many of the relationships and subtleties as possible, it would be best to feed the same content to the LLM all three ways.

    Most publicly available courseware modules are not consistently and explicitly designed in ways that would make this kind of slicing and dicing easy (or even possible). Luckily, I happen to know where can get my hands on some high-quality modules that are marked up in XML. Since I know just a little bit about XML and how these modules use it, I was able to have a conversation with ChatGPT about which XML to strip out, the pros and cons of converting the rest into English versus leaving them as XML, how to use the XML Document Type Definition (DTD) to teach the software about some of the explicit and implicit relationships among the module parts, and how to write the software that would do the work of converting the modules into input-output pairs. 

    By the end of the exploratory chat, it was clear that the work I want to accomplish requires more software programming skill than I have, even with ChatGPT’s help. But now I can estimate how much time I need from a programmer. I also know the level of skill the programmer needs. So I can estimate the cost of getting the work done. 

    To get this result, I had to draw on considerable prior knowledge. More importantly, I had to draw on significant language and critical thinking skills. 

    Anyone who ever said that a philosophy degree like mine isn’t practical can eat my dust. Socrates was a prompt engineer. Most Western philosophers engage in some form of chain-of-thought prompting as a way of structuring their arguments. 

    Skills and knowledge aren’t dead. Writing and thinking skills most certainly aren’t. Far from it. If you doubt me, ask ChatGPT, “How might teaching students about Socrates’ philosophy and method help them learn to become better prompt engineers?” See what it has to say. 

    (For this question, I used the GPT-4 setting that’s available on ChatGPT Plus.)

    Assessments aren’t dead either

    Think about how either of the projects I described above could be scaffolded as a project-based learning assignment. Students could have access to the same tools I had: an LLM like ChatGPT and an LLM-enhanced search tool like Bing Chat. The catch is that they’d have to use the ones provided for them by the school. In other words, they’d have to show their work. If you add a discussion forum and a few relevant tutorials around it, you’d have a really interesting learning experience. 

    This could work for writing too. My next personal project with ChatGPT is to turn an analysis paper I wrote for a client into a white paper (with their blessing, of course). I’ve already done the hard work. The analysis is mine. The argument structure and language style are mine. But I’ve been struggling with writer’s block. I’m going to try using ChatGPT to help me restructure it into the format I want and add some context for an external audience.

    Remember my earlier point about generative AI being a commoditizing force? It will absolutely commoditize generic writing. I’m OK with that, just as I’m OK with students using calculators in math and physics once they understand the math that the calculator is performing for them. 

    Students need to learn how to write generic prose for a simple reason. If they want to express themselves in extraordinary ways, whether through clever prompt engineering or beautiful art, they need to understand mechanics. The basics of generic writing are building blocks. The more subtle mechanics are part of the value that human writers can add to avoid being commoditized by generative AI. The differences between a comma, a semicolon, and an em-dash in expression are the kinds of fine-grained choices that expressive writers make. As are long sentences versus short ones, decisions about when and how often to use adjectives, choices between similar but not identical words, breaking paragraphs at the right place for clarity and emphasis, and so on. 

    For example, while I would use an LLM to help me convert a piece I’ve already written into a white paper, I can’t see myself using it to write a new blog post. The value in e-Literate lies in my ability to communicate novel ideas with precision and clarity. While I have no doubt that an LLM could imitate my sentence structures, I can’t see a way that it could offer me a shortcut for the kind of expressive thought work at the core of my professional craft.

    If we can harness LLMs to help students learn how to write…um…prosaic prose, then they can start using their LLM “calculators” in their communications “physics” classes. They can focus on their clarity of thought and truly excellent communication. We rarely get to teach this level of expressive excellence. Now maybe we can do it on a broader basis. 

    In their current state of evolution, LLMs are like 3D printers for knowledge work. They shift the human labor from execution to design. From making to creating. From knowing more answers to asking better questions. 

    We read countless stories about the threat of destruction to the labor force partly because our economy has needed the white-collar equivalent of early 20th-Century assembly line workers. People working full-time jobs writing tweets. Or updates of the same report. Or HR manuals. Therefore our education system is designed to train people for that work. 

    We assume that masses of people will become useless, as will education, because we have trouble imagining an education system that teaches people—all people from all socio-economic strata—to become better thinkers rather than simply better knowers and doers. 

    But I believe we can do it. The hard part is the imagining. We haven’t been trained at it. Maybe our kids will learn to be better at it than we are. If we teach them differently from how we were taught. 

    Likely short-term evolution of the technology

    Those of us who are not immersed in AI—including me—have been astonished at the rapid pace of change. I won’t pretend that I can see around corners. But certain short-term trends are already discernable to non-experts like me who are paying closer attention than we were two months ago. 

    First, generative AI models are already proliferating and showing hints of coming commoditization around the edges. We’ve been given the impression that these programs will always be so big and so expensive to run that only giant cloud companies will come to the table with new models. That the battle will be OpenAI/Microsoft versus Google. GPT-4 is rumored to have over a trillion nodes. That large of a model takes a lot of horsepower to build, train and run. 

    But researchers are already coming up with clever techniques to get impressive performance out of much smaller models. For example, Vicuña, a model developed by researchers at a few universities, is about 90% as good as GPT-4 by at least one test and has only 12 billion parameters. To put that in perspective, Vicuña can run on a decent laptop. The whole thing. Tt cost $300 to train (as opposed to the billions of dollars that have gone into ChatGPT and Google Bard). Vicuña is an early (though imperfect) example of the coming wave. Another LLM seems to pop up practically every week with new claims about being faster, smaller, smarter, cheaper, and more accurate. 

    A similar phenomenon is happening with image generation. Apple has quickly moved to provide software support for optimizing the open-source Stable Diffusion model on its hardware. You can now run an image generator program on your Macbook with decent performance. I’ve read speculation that the company will follow up with hardware acceleration on the next generation of its Apple Silicon microchips.

    “Socrates typing on a laptop” as interpreted by Stable Diffusion

    These models will not be equally good at all things. The corporate giants will continue to innovate and likely surprise us with new capabilities. Meanwhile, the smaller, cheaper, and open-source alternatives will be more than adequate for many tasks. Google has coined a lovely phrase: “model garden.” In the near term, there will be no one model to rule them all or even a duopoly of models. Instead, we will have many models, each of which is best suited for different purposes. 

    The kinds of educational use cases I described earlier in this post are relatively simple. It’s possible that we’ll see improvements in the ability to generate those types of learning content over the next 12 to 24 months, after which we may hit a point of diminishing returns. We may be running our education LLMs locally on our laptops (or even our phones) without having to rely on a big cloud provider running an expensive (and carbon-intensive) model. 

    One of the biggest obstacles to this growing diversity is not technological. It’s the training data. Questions regarding the use of copyrighted content to train these models are unresolved. Infringement lawsuits are popping up. It may turn out that the major short-term challenge to getting better LLMs in education may be access to reliable, well-structured training content that is unencumbered by copyright issues. 

    So much to think about…

    I find myself babbling a bit in this post. This trend has many, many angles to think about. For example. I’ve skipped over the plagiarism issue because so many articles have been written about it already. I’ve only touched lightly on the hallucination problem. To me, these are temporary obsessions that arise out of our struggle to understand what this technology is good for and how we will work and play and think and create in the future. 

    One of the fun parts about this moment is watching so many minds at work on the possibilities, including ideas that are bubbling up from classroom educators and aren’t getting a lot of attention. For a fun sampling of that creativity, check out The ABCs of ChatGPT for Learning by Devan Walton. 

    Do yourself a favor. Explore. Immerse yourself in it. We’ve landed on a new planet. Yes, we face dangers, some of which are unknown. Still. A new planet. And we’re on it.

    Strap on your helmet and go.

  • ChatGPT Wrote This Article and then Totally Stole My Job!

    ChatGPT Wrote This Article and then Totally Stole My Job!

    As I outlined recently in my “e-Literate’s Changing Themes for Changing Times” post, I am shifting my coverage somewhat. I’ll be developing and calling out tags I use for these themes so that you can go to an archive page on each one. This one will be listed under the “AI/ML” “third-wave EdTech,” and “future of work” tags.

    I’ve been fascinated by the rapid progression of ChatGPT article fads:

    1. Look at this weird thing that writes stuff!
    2. I asked ChatGPT a question—and here’s what it answered!!
    3. I asked ChatGPT to write this article—and it totally did!!!
    4. Students could use ChatGPT to student essays write essays!!!! End of the world or totally awesome?????
    5. I asked ChatGPT for suggestions about preventing students from using ChatGPT to cheat—and it gave me five great suggestions (and five terrible ones)!!!!!!

    Waitaminute. Let’s back up.

    Students finding easy ways to cheat is not exactly a new thing. Remember, “to Chegg” is a verb now. Let’s back up to fad #3. Writers are asking ChatGPT to write their articles, publishing those articles, and then advertising that the articles published under their by-line were written in 30 seconds by a machine.

    Do they want to get replaced by an algorithm?

    It seems to me we’re thinking about the problem that these algorithms present in the wrong way.

    At the moment, ChatGPT is a toy

    Language-generating algorithms ChatGPT and their image-generating cousins are toys in both good and bad ways. In a good way, they invite people to play. Anyone, whether or not they know anything about programming, can explore the possibilities and limitations of a certain kind of cutting-edge AI. It’s fun. And, like all good games, you learn by playing. You can learn what they can and can’t do well. How to write a good prompt. And so on. They make AI more accessible and, in a certain way, more understandable.

    They’re toys in the bad sense because they weren’t designed for any particular purpose and tend to do weird things for any given task. For example, I tried to get DALL-E 2—ChatGPT’s image-generating sibling—to create a retro comic book cover depicting ChatGPT as a scary vampire. I tried half a dozen different prompts. The best result I got was with “An evil vampire on the cover of a retro comic book called ‘ChatGPT’”. Here’s the best image it came up with:

    “An evil vampire on the cover of a retro comic book called ‘ChatGPT’” as interpreted by DALL-E 2

    The algorithm does evil comic-book vampires pretty well. Spelling a word that was explicitly spelled for it in the prompt? Not so much.

    Remember, a large language model algorithm like the one behind DALL-E 2 and ChatGPT don’t understand anything. It’s taking a bag of words and putting together pieces it found on the internet to construct something that seems like an appropriate response.

    ChatGPT, for example, is notoriously bad at math. Because it doesn’t calculate. That’s not what it is designed to do. Its developers have been working on improving this limitation. So I asked it about how good at math it is:

    Me: How good are you at math?

    ChatGPT: I am able to perform basic arithmetic operations, such as addition, subtraction, multiplication and division. My training data includes mathematical concepts and I have been trained to recognize and respond to arithmetic questions. However, it is important to note that I may make mistakes in complex or unconventional mathematical computations.

    Me What is 2794 + 83948?

    ChatGPT: 87,742

    That answer is incorrect.

    So ChatGPT, the miracle of modern technology destined to destroy all education, is not as “smart” as a calculator. Note that you can ask Google or your Apple search prompt the same question and reliably get a correct answer. OpenAI (the makers of OpenAI) are likely trying to combine much simpler but mathematically reliable algorithms with the GPT-3 algorithm that underlies ChatGPT and DALL-E 2. The fact that they haven’t solved this problem yet suggests it is non-trivial. These different types of algorithms can be challenging to mesh. Many of the more impressive applications of artificial intelligence and machine learning are made possible by an ensemble of different algorithms working together, some of which fit together more naturally than others.

    Don’t get me wrong; the math problem will be solved. The OpenAI engineers will crack it just as they are now pretty good at preventing ChatGPT from making stock predictions or answering questions about current events when its data set hasn’t been updated since 2021. But it’s a game of whack-a-mole. Because you can ask ChatGPT anything, people do ask it anything. The creators are learning a lot about the questions people ask and what can go wrong with the answers. This new knowledge will help them design more specific solutions. But a general-purpose prompt tool like ChatGPT will be hard to make good at solving any one particular problem.

    I’m not convinced that ChatGPT, as it exists today, represents a big leap forward in essay cheating. It has length limitations, has to be fact-checked, can’t produce references, and spits out highly variable quality of reasoning and argumentation. Students would learn more by trying to fix the problems with a ChatGPT-generated draft than they would by going to a traditional essay mill.

    Short answer questions are a different matter. ChatGPT is already dangerous in this area. But again, students can already “Chegg” those.

    Yes, but…

    Could somebody write a better program specifically for writing school essays? Or magazine articles? Yes. That work is already underway.

    So what do we do about the essay cheating problem? Let’s start with the two most common answers. We can develop algorithms that detect prose that was written by other algorithms. That too is already underway. So we’ll have yet another flavor of the cheating/anti-cheating arms race that benefits nobody except the arms dealers. The anti-cheating tools may be necessary as one element of a holistic strategy, but they are not the ultimate answer.

    Second, we can develop essay-writing prompts and processes that are hard for the algorithms to respond to. This would be useful, partly because it would be good for educators to rethink their stale old assignments and teaching practices anyway. But it’s a lot of often uncompensated work for which the educators have not been trained. And it ends up being another arms race because the algorithms will keep changing.

    We miss the point if we respond to language-generating AI as a static threat that might become more sophisticated over time but won’t fundamentally change. ChatGPT is just a friendly way for us to develop intuitions about how one family of these algorithms works at the moment. You’re wrong if you think it is a one-time shock to the system. We’re just at the beginning. The pace of AI progress is accelerating. It is not just going to get incrementally better. It is going to radically change in capabilities at a rapid pace. It will continue to have limitations, but they will be different limitations.

    So what do we do?

    How about talking to the students?

    When adaptive learning hit peak hype, a glib response to teacher hysteria started making the rounds: “If you [teachers] can be replaced by a computer, then you probably should be.”

    Doesn’t that apply…um…generally?

    If all students learn is how to use ChatGPT to write their essays, why wouldn’t their hypothetical future employer use ChatGPT instead of hiring them? Why would students spend $30K, $40K, $50K, or more a year to practice demonstrating that a free-to-use piece of software does their best work for them? Students need to learn the work these tools can do so they can also understand the work the tools can’t do. Because that is the work the students could get paid for. Technology will make some jobs obsolete, leave others untouched, change some, and create new ones. These categories will continue to evolve for the foreseeable future.

    At a time when students are more conscious than ever about the price-to-value of a college education, they ought to be open to the argument that they will only make a decent living at jobs they can do better than the machine. So they should learn those skills. Why learn to write better? So you can learn to think more creatively and communicate that creativity precisely. Those are skills where the primates still have the advantage.

    Once we engage students openly and honestly on that point, we will start building a social contract that will discourage cheating and establish the foundational understanding we need for rethinking the curriculum—not just to keep from falling too far behind the tech but to help students get out in front of it. The current limitations of these AI toys demonstrate both the dangers and the potential. Suppose you want to apply the technology to any particular domain. In that case, whether it’s math, writing advertising copy, or something else, you need to understand how the software works and how the human expertise and social or business processes work. Whole echelons of new careers will be created to solve these problems. We will need thinkers who can communicate. Learning how to formulate one’s own thoughts in writing is an excellent way to learn both skills.

    Fighting the tech won’t solve the problem or even prevent it from getting worse. Neither will ignoring it. We have to engage with it. And by “we,” I include the students. After all, it’s their futures at risk here.

    (Disclaimer: This blog post was written by ChatGPT.)

    (I’m kidding, of course.)

    (I am able to perform basic humor operations, such as generating dirty limericks and “your momma is so ugly” jokes. My training data includes humorous concepts, and I have been trained to recognize and respond to knock-knock questions. However, it is important to note that I may make mistakes in complex or unconventional humor.)

  • I Would Have Cheated in College Using ChatGPT

    I Would Have Cheated in College Using ChatGPT

    As I outlined recently in my “e-Literate’s Changing Themes for Changing Times” post, I am shifting my coverage somewhat. I’ll be developing and calling out tags I use for these themes so that you can go to an archive page on each one. This one will be listed under the “AI/ML” “third-wave EdTech,” and “future of work” tags.

    ChatGPT is creating all kinds of buzz about students cheating on essays. So it got me thinking. If I had had access to a tool like ChatGPT when I was in college, would I have used it to cheat?

    A robot writing an essay
    “A robot writing an essay” as interpreted by DALL-E 2

    Yes. Absolutely. One hundred percent. But I wouldn’t have thought of it as cheating.

    Cheating is a state of mind

    When I was in college, I had a massive chip on my shoulder. If I caught the slightest whiff that the professor didn’t care whether I was learning or the assignment was not well-designed to help me learn something, I would immediately flip into “grudge” mode.

    I never cheated. The whole point, in my immature mind, was to prove that I was smarter than the professor. So I would stay up late, party the night before the assignment was due, and give myself three or four hours to write it the next morning, hung over, before I had to turn it in. To give you a sense of just how far these grudges went, I didn’t actually enjoy partying very much. Sometimes I would go out of my way to do so because of the stupid assignment. It was all part of a game to challenge myself. No pain, no game.

    I wasn’t always so self-destructive about it. If I felt the professor was genuinely interested in my learning but had simply written an assignment I wouldn’t learn from, I’d take the writing prompt and try to be as creative—and subversive—as possible. But that only worked when the professor would understand and appreciate the joke. If I had a professor who had (in my judgment) given me a coloring book and was going to grade me on whether I colored inside the lines, that triggered my worst adolescent self. While I didn’t cheat in the conventional sense, my goal was to minimize effort required to get an adequate grade and show myself how smart I was in the process. My social contract with the teacher was no less broken than the student who copied somebody else’s work. I once defined cheating as “engaging in behaviors that are intended to facilitate passing without learning.” By that standard, I cheated my ass off.

    Something like ChatGPT would have been part of the game to me. If you’ve played around with it at all (or read some articles about it), you’ll know that writing prompts that lead the AI to generate good text is something of an art. I would have spent a lot of time crafting the best prompt possible. Then I would have edited the output. I still would have wanted to produce a good essay. If this process took more time to produce the end result than just writing it from scratch myself, it wouldn’t have mattered to me. I would have subverted the assignment into something that actually challenged me while flipping the bird to the instructor who had the temerity to underestimate or underappreciate their students in general and me in particular. That was the whole point of the game. I wanted to learn. And I wanted to care. Nothing pissed me off more than a professor who wasted my educational time.

    ChatGPT wouldn’t have violated my “no cheating” rule because I wouldn’t have been cheating according to my rules.

    Faculty tend to think that cheating is defined by their rules and the college honor code. The reality is far more complex. For me, it was heavily influenced by my social contract with each teacher, whether I felt they were holding up their end of the bargain, and how I could turn every assignment into a game that was a fun mental challenge. Other students may be influenced by which assignments they think are important for their career goals, how much work is reasonable to expect of them and, very often, or whether they think the instructor cares about their learning.

    I can’t emphasize that last point enough. I’ve conducted a fair few focus groups with students over the years. The results consistently supported the research evidence that students are heavily influenced by whether they believe their teacher cares about their learning. And this manifests in surprising ways. I remember one particular focus group vividly. We were talking about what factors cause students to engage in with a class more than they expected or planned to. They all agreed that having a teacher that cared about their learning was a major factor. I asked them how they can tell if a teacher cares. I got some of the usual, expected answers like “she knows my name” or “he talks to me after class to see how I’m doing.” One surprising answer that elicited a lot of nods from the group was, “She picked a random student during class and asks them a question about the reading. She didn’t grade us on the answer but it’s embarrassing if you don’t know it.”

    “You liked that?” I asked, somewhat incredulously.

    “No,” he replied. “I hated it. But I know she did it because she wanted us to learn. So I did the readings.”

    ChatGPT as a tutor

    The current fad with journalism about ChatGPT is to ask the bot some version of, “If you were writing this article, what would you write?” The value of that exercise depends on the incisiveness of the commentary regarding the output, the prompt, and the relationship to writing craft. Examples of this genre can range from pretty interesting to pretty terrible, depending on the quality of thinking of the human “writer” involved. Bad articles end either in “WHEEEEE!” or “NO HUMAN WILL EVER LEARN HOW TO WRITE A COMPLETE SENTENCE AGAIN!” Better articles give us insight into what the tool is good for as a writing aid and how that reflects on writing craft.

    Let’s see how well I fare at this genre. Can ChatGPT write this blog post for me? The base question I’m after is how to use a tool like ChatGPT to engage students in the craft of writing the way it would have engaged me. As you’ll learn by playing around with the tool or even reading some of the better articles on it, part of the craft in getting useful answers out of ChatGPT is writing a good prompt. My prompts in the experiments below are in bold. ChatGPT’s answers follow.

    I’m a college philosophy professor. I have created an assignment asking students to write an essay on a particular topic. They are allowed to use ChatGPT to generate their first draft but will be graded on the accuracy, originality of thought, and writing quality of their final draft. They will also be evaluated on the quality of the ChatGPT prompt that they write. What are the educational strengths and weaknesses of this assignment?

    There are both strengths and weaknesses to using a tool like ChatGPT for a college philosophy assignment. One potential strength is that it can help students generate ideas and get started on their essays. This can be particularly helpful for students who may be struggling to come up with a topic or approach to the assignment.

    Another potential strength is that using ChatGPT can help students practice using technology in their writing. This is an important skill in today’s world, and it can help students develop the ability to effectively use tools like ChatGPT in their future academic and professional endeavors.

    On the other hand, there are some potential weaknesses to this assignment. One is that ChatGPT is a tool that generates text based on prompts, but it does not have the ability to understand the meaning or significance of the words it generates. This means that the essays produced using ChatGPT may not be as original or thought-provoking as those produced by the students themselves.

    Another potential weakness is that using ChatGPT could discourage students from engaging with the material in a deep and meaningful way. Instead of reading and thinking critically about the assigned texts, they may simply use ChatGPT to generate a first draft and then submit it without further thought or analysis.

    Overall, the use of ChatGPT in a college philosophy assignment has both strengths and weaknesses. It can be a helpful tool for generating ideas and getting started on an essay, but it should not be used as a substitute for careful reading, thinking, and writing.

    My dialogue with ChatGPT

    Meh. ChatGPT’s answer has all the hallmarks of an adequate but unimaginative high school essay, right down to the form. It has not produced an essay that is “original or thought-provoking” or written content that engages with the prompt “in a deep and meaningful way.”

    We can marvel at the ability for the technology to do this well. It is genuinely miraculous and merits a brief pause for a moment of wonder. That said, the content itself is not enlightening or even particularly interesting by human writing standards.

    Let’s try again. I’ve altered the prompt using language from ChatGPT’s output in our first try. To make the changes more obvious, I’m showing the text I deleted from the original prompt and underlining the text I added:

    I’m a college philosophy professor. I have created an assignment asking students to write an essay on a particular topic. They are allowed to use ChatGPT to generate their first draft but will be graded on the accuracy, originality of thought, and writing quality of their final draft. They will also be evaluated on the quality of the ChatGPT prompt that they write. What are the educational strengths and weaknesses of this assignment? How can I improve this assignment to encourage students to produce more original, thought-provoking assignments in which the students engage with the material in a deep and meaningful way?

    One way to improve this assignment would be to provide more guidance on how to use ChatGPT effectively. For example, you could provide students with examples of good and bad ChatGPT prompts, and explain how to write a prompt that will generate thoughtful, original responses from the tool.

    Another way to improve the assignment would be to require students to engage more deeply with the material before using ChatGPT. For example, you could ask them to read and annotate the assigned texts, or to participate in a class discussion, before using ChatGPT to generate their first draft. This will help ensure that they have a strong understanding of the material and are able to generate original, thought-provoking ideas on their own.

    In addition, you could also consider modifying the grading criteria for the assignment. For example, you could place less emphasis on the accuracy of the ChatGPT-generated text, and instead focus on the student’s own writing and analysis in the final draft. This would encourage students to take the time to engage with the material and produce original, well-written essays.

    Overall, there are several ways you could improve this assignment to encourage students to produce more original, thought-provoking work. By providing more guidance on how to use ChatGPT effectively, requiring deeper engagement with the material, and modifying the grading criteria, you can help students develop the skills and habits needed to produce high-quality work in their philosophy studies.

    My dialogue with ChatGPT

    This answer is more interesting. The first suggestion—teach students how to write effective prompts—is a good start. I do think this is a skill that will be taught widely in schools within five years. (It should be sooner than that and it realistically might be later.) I don’t agree with the suggestion to de-emphasize the accuracy of ChatGPT’s text. Evaluating the quality of source material is an essential skill which becomes more important, complex, and interesting with the introduction of tools like this one. That said, a lot depends on context. I could imagine assignments for which I would take ChatGPT’s advice on this point. Notice, though, that the structure of the “essay” ChatGPT generated is still pretty stiff and formulaic. Leaning too heavily on a tool like this could have the effect of cultivating competent writing at the expense of stifling great writing.

    Still, I can see the shape of a pedagogical process—and preferably a supporting end-to-end tool—that teaches many of the skills involved with good writing, including some hard ones like checking sources and editing—while including some elements of creativity. If it is scaffolded properly—again, with the right tool and process but also with a good, solid rubric—it could enable educators to spend more of their time honing in on specific aspects of the writing process with less drudgery. Particularly if used judiciously as part of the writing curriculum rather than the whole thing, it could be quite useful.

    It’s also honest. It strengthens rather than weakens the social contract between student and educator by allowing the students to use a tool as long as they are open about it and are using it as part of a genuine learning process rather than a shortcut around thought work.

    Would I use ChatGPT to help me write blog posts?

    In principle, I have no problem with the idea of using machine-generated text in e-Literate posts as long as it is properly attributed. In practice, I haven’t been able to figure out a way to make it useful.

    Part of the value of e-Literate is that it can be surprising in both content and form. Novelty teaches while entertaining. Could a tool like ChatGPT develop the right sort of novelty to fit with this blog? On the surface, it probably could. We’re already starting to see examples of the tool being asked to write an article on X subject in the style of Y person. If trained on the thousands of posts I’ve written, I suspect that a tool like ChatGPT, or maybe the next generation of it, could learn to use more em dashes, write convoluted sentences, and be more snarky. It might even incorporate some themes that show up in my posts.

    But it doesn’t actually understand anything that it writes. ChatGPT distills and synthesizes answers that have already been given. It can only write about ideas that I have already thought of and written about. It can’t write about the idea I’m going to have tomorrow. e-Literate isn’t about what I know. It’s about what I’m learning. As such, I don’t yet see how I could get much value out of a tool like ChatGPT in the foreseeable future, at least for this blog. The only exception I can think of is for posts like this one that are about the tool. ChatGPT didn’t really write part of this post for me. It generated artifacts for me to analyze in my own writing.

    These lines of demarcation—the lines between when a tool can do all of a job, some of it, or none of it—are both constantly moving and critical to watch. Because they define knowledge work and point to the future of work. We need to be teaching people how to do the kinds of knowledge work that computers can’t do well and are not likely to be able to do well in the near future. Much has been written about the economic implications to the AI revolution, some of which are problematic for the employment market. But we can put too much emphasis on that part. Learning about artificial intelligence can be a means for exploring, appreciating, and refining natural intelligence. These tools are fun. I learn from using them. Those two statements are connected.

    Would I teach writing using ChatGPT?

    If I were teaching writing today, would I use an AI tool? In practice, probably not, simply because it would be too much work to cobble together the pieces. ((I have now guaranteed that I will get an avalanche of emails from EdTech startups claiming to solve this problem. Please direct your messages to my chatbot.)) This is the perennial challenge of EdTech, which, on balance, creates a vastly underestimated drag on the amount of time educators have to put into the thought work of delivering high-quality education. In principle, though, yes, I absolutely would. I have a clear picture in my head of what I would need in terms of the EdTech and what sorts of writing work I would (and wouldn’t) use it for.

    Do I think higher ed is ready for widespread adoption of these tools? That’s a harder question. Teaching this way requires a new skillset. Higher ed has abysmally under-resourced professional development support for teaching, on the whole. Also, teaching this way isn’t what our PhD system trains young academics to aspire to. Many will see it as a dumbing down of the work they have dedicated their lives to. And if implemented poorly, it can easily turn into that.

    So will AI text generation tools revolutionize or kill college writing? Both! Neither! For sure! Probably! Eventually! Somewhat! It’s…complicated.

    As usual.

  • Seeing the Future: Developing Intuitions About Artificial Intelligence

    Seeing the Future: Developing Intuitions About Artificial Intelligence

    I know I’ve been on a bit of a tear lately about artificial intelligence (AI). I promise e-Literate won’t turn into the “all AI all the time” blog. That said, since I have identified it as a potential factor in a coming tipping point for education, I think it’s important that we sharpen our intuitions about what we can and can’t expect to get from it.

    Plus, it’s fun.

    In a recent post, I quoted an interview with experts in the field who were talking about playing with AI tools that can generate images from text descriptions as a way of expressing their creativity. And in my last post, I included an image from one such tool, DALL-E 2, created from the prompt “A copy of the sculpture “The Thinker” made by a third grader using clay.”

    A copy of the sculpture “The Thinker” made by a third grader using clay as interpreted by DALL-E 2

    In this post, I will use this image and the tool it created as a jumping-off point for exploring the promise and limitations of large cutting-edge AI models.

    Interpreting art

    To say that I lack well-developed visual skills would be an understatement. When I’m thinking, which is generally whenever I’m awake, I am usually looking at the inside of my skull. I’ve been known to walk into fire hydrants and street signs with some regularity.

    My wife, on the other hand, has an eye. She took private sculpture lessons in high school from Stanley Bleifeld. She did a brief stint at art school before turning to English. She teaches our grandchildren art. And she loves Rodin, the sculptor who created the thinker. I picked the image for the post mainly because I liked it. Her reaction to it was, “That looks nothing at all like the original. Where are the hunched shoulders? The crossed elbow? Where’s the tension in the figure? And what’s with the hair? That’s not what a third grader would make.”

    The Thinker
    The Thinker by Auguste Rodin is licensed under CC-CC0 1.0

    So we looked at other options. DALL-E 2 generates four options for each prompt, which you can further play with. Here are the four options that my prompt generated:

    The bottom one is the one she thought best captured the original and is most like what a third grader would produce.

    The model did well with “clay.” I tested its understanding of materials by asking it to copy the famous sculpture in Jell-O. In all four cases, it captured what Jell-O looks like very well. Here’s the best image I got:

    A copy of the sculpture “The Thinker” made in green Jell-O as interpreted by DALL-E 2.

    The AI clearly knows what green Jell-O looks like, down to the different shades that can come from light and food coloring. (The Jello-O mold as the seat is a nice touch.) That’s not surprising. The AI likely had many examples of well-labeled images of Jello-O on the internet.

    It struggled with two aspects of the problem I gave it to solve. First, what are the salient features of The Thinker in terms of its artistic merit? Which details matter the most? And second, how would artists at different ages and developmental stages see and capture those features?

    Let’s look at each in turn.

    Artistic detail

    My wife has already given us a pretty good list of some salient features of the sculpture. The subject is literally and figuratively pensive. (Puns intended.) Can we get the AI to capture the art in work? My first experiment was to try asking it to interpret the sculpture through the lens of another artist. So, for example, here’s what I got when I asked it to show me a painting of the sculpture by Van Gogh:

    Painting of the sculpture The Thinker by Vincent Van Gogh as interpreted by DALL-E 2.

    Interesting. It gets some of the tension and some of the balance between detail and lack of detail (although that balance is also consistent with Impressionist painting). But all four of the thinker images I got back for this prompt had Van Gogh’s head on them. This is probably because Van Gogh’s famous portraiture is self-portraiture. What if we tried a renowned portrait artist like Rembrandt?

    Painting of the sculpture The Thinker by Rembrandt as interpreted by DALL-E 2.

    I’m not sure I would describe this figure as pensive, exactly. To my (poor) eye, the tension isn’t there. Also, all four examples came back with the same white hat and ruffle. The AI has fixated on those details as essential to Rembrandt’s portraits.

    What if we stretched the model a bit by trying a less conventional artist? Here’s an example using Salvador Dalí as the artist:

    Painting of the sculpture The Thinker by Salvador Dalí as interpreted by DALL-E 2.

    Hmm. I’ll leave it to more visual folks to comment on this one. It doesn’t help me. I will note that all four images the AI gave me had that strange tail coming out of the back of the head. It has made a generalization about Dalí’s portraiture.

    I won’t show you DALL-E 2’s interpretation of Hieronymus Bosch’s version of The Thinker, not because it’s gross but because it just didn’t work at all.

    DALL-E 2 is a language model tacked onto an image model. It’s interpreting the words of the prompt based on analyzing a large corpus of text (e.g., the internet) and mashing that up with visual features it’s learned from analyzing a large corpus of images (e.g., the internet). But the connection between the two is loose. For example, even though it’s probably digested many descriptions and analyses of The Thinker, it doesn’t translate that information to the visual model. My guess is that if I built a chatbot using the underlying GPT-3 language model and asked it about the features that are considered important in The Thinker as a work of art, it could tell me. DALL-E 2 doesn’t translate that information about salient features into images.

    How could you fix this if you wanted to build an application that can visually re-interpret works of art while preserving the essential features of the original? I’m going to speculate here because this gets beyond my competence. These models can be tuned by training them on special corpi of information. I’m told they’re not easy to tune; they’re so complex that their behavior can be unpredictable. But, for example, you could try to amplify the art history analyses in the information that gets sent from the language model to the visual model. I’m not sure how one would get the salient features picked up by the former to be interpreted by the latter. Maybe it could elaborate on your prompt to include details that you didn’t. I don’t know. Remember my post about the miracle, the grind, and the wall in AI? This would be the grind. It would be a lot of hard work.

    Artistic development

    Getting the salient details of art is hard enough. But I also asked it to interpret those details not through the eyes of a specific artist but through the eyes of a person at a particular developmental level. A third grader. GPT-3 does have a model of sorts for this. If you ask it to give answers that are appropriate for a tenth grader, you will get a different result than if you ask it to respond to a first-year college student. Much of the content it was trained on undoubtedly was labeled for grade level and/or reading level. It doesn’t “know” how tenth graders think but it’s seen a lot of text that it “knows” were written for tenth graders. It can imitate that. But how does it translate that into artistic development?

    Here’s what I got when I asked DALL-E 2 to show me clay copies of The Thinker created by “an artistic eighth grader”:

    These are, on the whole, worse. We want to see sculptures that more accurately capture the artistically salient features of the original. Instead, we get more hair and more paint.

    How could you get the model to capture artistic development? Again, I’ll speculate as a layperson. The root of the problem may well be in the training data. The internet has many, many images. But it doesn’t have a large and well-labeled set of images showing the same source image (like the Rodin sculpture) being copied by students at different age levels using different media (e.g., clay, watercolors, etc.). If so, then we’ve hit the wall. Generating that set of training data may very well be out of reach for the software developers.

    Language is weird

    I’ll throw one more example in just for fun. By this point in my experiment, I had gotten bored with The Thinker and was just messing around. I asked the AI to show me “a watercolor painting of Harry Potter in a public restroom.” Now, there are two ways of parsing this sentence. I could have been asking for “(a watercolor painting of Harry Potter) in a public restroom” or “a watercolor painting of (Harry Potter in a public restroom)”. DALL-E 2 couldn’t decide which one I was asking for, so it gave me both in one image:

    A watercolor painting of Harry Potter in a public restroom as interpreted by DALL-E 2

    These models are tricky to work with because it’s very easy for us to wander into territory where we’re recruiting multiple complex aspects of human cognition such as visual processing, language processing, and domain knowledge. It’s hard to anticipate the edge cases. That’s why most practical AI tools today do not use free-form prompts. Instead, they limit the user’s choices through the user interface to ensure that the request is one that the AI has a reasonable chance of responding to in a useful way. And even then, it’s tricky stuff.

    If you’re a layperson interested in exploring these big language models more, I recommend reading Janelle Shane’s AI Weirdness blog and her book You Look Like a Thing and I Love You: How Artificial Intelligence Works and Why It’s Making the World a Weirder Place.

  • AI/ML as Copilot

    AI/ML as Copilot

    I will stick with the topic of artificial intelligence and machine learning (AI/ML) for today’s post because I keep getting feedback suggesting there’s a lot of interest in it. Since I’m on a bit of a streak, I feel compelled to make some caveats before jumping in.

    First, I am not a software engineer or an AI/ML expert. In fact, I’m not an expert in many of the topics I write about. I just happen to be pretty good at making inferences from a small amount of information and understanding. I write about what I’m learning rather than what I know. For many years, the tagline for this blog was “What I’m learning about online learning.” ((“What I’m learning about stuff that interests me and is in some way relevant to digitally-enabled learning” seemed too long.)) Since this blog is about learning rather than knowing, the corollary is that I invite you to educate me if you know something I don’t. Please tell me if I’m wrong, if I’m missing something, or even if you think I’m right.

    Second, I’ll be writing again about the big new AI language models that have been taking the world by storm lately. On the one hand, I risk adding to misperceptions with this focus. A lot of folks are writing about these models now because they’re so sexy, surprising and, frankly, potentially dangerous. In reality, AI/ML is a large, diverse, and ever-growing family of computational techniques, many of which have very different characteristics from each other. On the other hand, the big models are useful to write about precisely because they are on the extreme end of the spectrum regarding their alienness. They highlight some problems that may be more subtle and harder to see in other techniques.

    My last caveat is that I will be responding to an interview conducted by the great Ben Thompson of Stratechery. Specifically, I’ll be quoting from one of his subscription-only articles. Since this is his bread and butter, I’m mindful of putting too much of his paid content on the public internet. Luckily, it’s a very long article and I’ll only be quoting a small fraction of it. While I think it likely meets the criteria of fair use, more importantly, I’m hoping Ben will see it as an advertisement. I’m a fan. I don’t pay for many newsletters. I do pay for his. If you want to understand the intersection of tech and business, you should too. He’s a fantastic writer and an original thinker.

    This post happens to be an interview, which is unusual for Stratechery. The interviewees are Daniel Gross and Nat Friedman. Ben identifies them first as VCs, but their salient credentials are that they both worked extensively in tech and are real experts in AI/ML. His post is fascinating in its entirety. I’m going to focus on a few aspects that are salient for EdTech and that resonate with my recent screeds on AI/ML.

    Spooky and kooky

    In my post about using GPT-3 to create a philosophy tutor chatbot, I wrote about the miracle, the grind, and the wall. First, these models do something that blows your mind. That gets you excited. As you try to turn that moment of exhilaration into a reliable and scalable piece of software, you discover that reliable and scalable are both arduous work. Eventually, you hit a wall you can’t get past. And it’s hard to predict in advance where that wall will be.

    Nat Friedman oversaw the creation of GitHub Copilot, which uses a version of GPT-3 to suggest code to developers in real-time as they are writing software. Here’s what he said about what it was like:

    he thing I would always say with those models is that they alternate between spooky and kooky. So half the time or some fraction of the time, they’re so good, it’s spooky like, “How did it figure that out? It’s incredible. It’s reading my mind,” or “It knows this code better than I do.” Then sometimes it’s kooky, it’s just so wrong, it’s nonsense, it’s ridiculous. So when it was wrong, it was really wrong. It turned out from testing it in the Q&A scenario that when you actually asked the thing a question and it gave you more often than not a wrong answer, you got very irritated by it — this was an extremely bad interaction. So we knew that it couldn’t be some explicit Q&A interaction. It couldn’t be something where you ask a question and then 70 percent of the time you get a useless answer. It had to be some product where it was serving you suggestions when it has high confidence, but it wasn’t something you were asking for and then getting disappointed by….

    [I]t turns out in retrospect, we know this now and we didn’t know it at the time, the question that we were trying to answer was, “How do you take a model which is actually pretty frequently wrong and still make that useful”? So you need to develop a UI which allows the user to get a sense and intuition themselves for when to pay attention to the suggestions and when not to, and to be able to automatically notice, “Oh, this is probably good. I’m writing boilerplate code,” or “I don’t know this API very well. It probably knows it better than I do,” and to just ignore it the rest of the time.

    Stratechery

    Friedman’s comments highlight that even Microsoft, using one of the most advanced AI models on the planet, could only get a useful answer from the AI about 30% of the time after carefully training it on the vast body of software code in GitHub that had been tested and validated as working.

    There isn’t even a moment’s consideration given to having the model replace the programmer. It’s wrong 70% of the time. That might not be true always and forever but it’s true now with an army of skilled engineers using one of the best models available. The product is called Copilot. Not only is there a human in the loop; the human is in charge. A lot of thought went into designing the software so that expert humans will feel comfortable ignoring it and not annoyed that it’s wrong so often:

    So it’s funny because a lot of the ideas we had about AI previously were this idea of dialogue. The AI is this agent on the other side of the table, you’re thinking about the task you want to do, you’re formulating it into a question, you’re asking, and you’re getting a response, you’re in dialogue with it. The Copilot idea is the opposite. There’s a little robot sitting on your shoulder, you’re on the same side of the table, you’re looking at the same thing, and when it can it’s trying to help out automatically. That turned out to be the right user interface….

    So from the June realization that we should do something, I think it was end of summer, maybe early-September by the time we concluded chatbots weren’t it. Then it really wasn’t until February of the next year that we had the head exploding moment when we realized this is a product, this is exactly how it should work…. So now, it’s very obvious. It seems like the most obvious product and a way to build, but at the time, lots of smart people were wandering in the dark looking for the answer.

    Stratechery

    This is not the way a lot of EdTech is designed. The industry tends to develop models that automate various aspects of course design, teaching, or support and shoot it straight to the students without much thought about how those students will handle or even recognize, errors. Even when there’s a human in the loop, we’re not typically designing these products for expert humans. Instead, we’re using the humans to review—which might mean spot check—the machine-generated output or tool, which is intended for use by explicitly non-expert humans.

    “A Robot Tutor in the Sky,” by the DALL-E 2 algorithm

    The best systems using tech that is extremely useful but highly unreliable are designed to make the role of human judgment obvious to see and easy to apply. As Friedman notes, getting this right is very hard.

    Getting the interaction model right

    Thompson’s response to Friedman was characteristically insightful:

    The reason is that people want to anthropomorphize everything and they want to put everything in human terms. The whole point of a computer is it just operates utterly and completely different than humans do. At the end of the day, it’s still calculating ones and zeros. So everything has to be distilled to that and it just does it at tremendously fast speed, unimaginable speed, but that is so completely different than the way that a human mind works that that’s how whatever was kooky or spooky I’m sure was completely and utterly logical to the computer. It strikes me that this is why the chat interface was wrong, because what it was doing was it was taking this intelligence, and it was actually accentuating the extent to which it was different than humans by trying to put it as a human, as if you’re talking to someone, and it was actually essential to come up with a completely different interface that acknowledged and celebrated the fact that this intelligence actually functions completely and utterly differently than humans do.

    Stratechery

    The three lessons here are (1) make the AI visible, (2) make it clear that the AI is not some simulacrum of human intelligence but rather is a tool that works imperfectly, and (3) create an experience that encourages users to exercise judgment when evaluating the output of the AI.

    Here’s Friedman again about Copilot:

    The thing that Copilot gave us that we, again, only realized in retrospect was this randomized psychological reward. It’s like a slot machine where the ongoing cost of using it at any given moment is not very high, but then periodically you hit this jackpot where it generates a whole function for you and you’re utterly delighted. You can’t believe it, it just saved you 25 minutes of Googling and Stack Overflow and testing. That happens at random intervals, so you’re ready for the next randomized reward, it has this addictive quality as a result. Whereas people frequently have ideas that are like, “Oh, the agent is going to write a huge pull request for you and it’s going to write a huge set of changes across your code, you’re going to review that.”

    Stratechery

    Copilot turns a bug into a feature. The unreliability leaves you delighted when the product gives you something useful rather than angry when it doesn’t. This only works because the user explicitly understands that the product is 70% unreliable but finds it helpful—and delightful—anyway.

    Again, that’s not how EdTech typically is designed and it’s certainly not how it’s usually sold. We often present the tech to non-expert users as a virtual helper that’s positioned as a tutor or advisor.

    It doesn’t have to be that way. First, we should spend more energy in EdTech using AI/ML to help expert users—i.e., the educators—more efficiently and effectively leverage their expertise to help the non-expert users—i.e., the students. Second, we can create interfaces for non-expert users that turn the unreliability of AI models from a bug into a feature. One portion of the interview turns to new models that enable users to generate novel images from text. Friedman talks about people using one such model, called Midjourney, that they access through a discussion forum called Discord:

    if you ever watch a really creative person sit over their shoulder and watch them use Midjourney for an hour, you find that what they’re doing is not one text-to-image. They’re writing a prompt, they’re generating a bunch of images, they’re generating variance of those images, they’re remixing ideas, they might be riffing off someone else in the Discord channel, they’re exploring a space, you’re exploring latent space in a way, and then pinning the elements of it that you like, and you’re using that for creativity and ideas, but also to zero in on an artifact or an output that you’re trying to produce, and those are different modes.

    Stratechery

    This was similar to my experience trying to write a philosophy tutor. It was creative and fun. It felt a bit like teaching a student with strengths and limitations that I was trying to understand. I would probe, adjust, and probe again. I loved it. And when I had reached my limit with the chatbot, yes, I was tempted to throw it out and start fresh with another idea.

    Friedman, noting the popularity of Midjourney on Discord, described an example user he heard about from the company’s founder:

    There was one David was telling me about recently who’s a trucker, who when he stops at truck stops, he canceled his Netflix, and now what he does is he just makes images for a couple hours before bed, and he’s utterly transfixed by this. To me, that seems like it’s just objectively better than watching Netflix and binging a show; it’s exploring the space of your own ideas and creativity and seeing them fed back to you. So it turns out there’s a lot of people who have this creative impulse and just didn’t have the tools, the manual skills to express it and to create art, and something like Midjourney or something like Stable Diffusion gives them that, and that’s incredibly exciting.

    Stratechery

    In education, we discuss the need to expose our students to AI/ML and develop some literacy. This is equally true in the workforce, by the way. And yet, when we do expose students to the tech, make it as we tend to make it invisible as possible.

    The bottom line

    Some readers have interpreted me as an AI/ML skeptic, at least as applied to EdTech. The opposite is true. I’m enthusiastic about its potential and some of the current uses I’ve seen. I simply believe that we are immature in our thinking about how to apply it constructively to our domain. While we’re hardly alone in that, we have an extra responsibility of care as educators. We should be looking for new models that previously weren’t possible rather than just trying to automate and accelerate old models. In doing so, we should embrace the limitations of the technology and let them inspire fresh thinking.

  • AI/ML in EdTech: The Miracle, The Grind, and the Wall

    AI/ML in EdTech: The Miracle, The Grind, and the Wall

    I had a conversation with a friend last night about the counter-intuitive challenges of working with AI/ML and the implications for using them in EdTech. I decided it might be useful to share my thinking more broadly in a blog post.

    Essentially, I see three stages in working with artificial intelligence and machine learning (AI/ML). I call them the miracle, the grind, and the wall. These stages can have implications for both how we can get seduced by these technologies and how we can get bitten by them. The ethical implications are important.

    The Miracle

    One challenge with AI/ML is how deceptively easy it is to produce mind-blowing demos with these tools. For example, I spent some time playing with GPT-3 as a learning exercise. GPT-3 is one of several gigantic AI models that can do some pretty miraculous things with natural language. (Google has the other prominent gigantic model.) One reason I started with GPT-3 is that it can be programmed using natural language. For example, you can tell it, “You are a helpful chatbot that teaches first-year college students about philosophy” and voila! You have a helpful philosophy-teaching chatbot.

    It’s not quite that simple. For example, I found that GPT-3’s idea of what a first-year college student understands about philosophy differed from mine. I got better results when I asked it to target 11th grade. I couldn’t have known that in advance. GPT-3 is a neural network of 175 billion parameters. It has indexed large swathes of the internet and many books. But it doesn’t store all that information, exactly. It distills it in very complex ways. In fact, GPT-3 and similar models are so complex that even the programmers who made them can’t explain why they produce specific responses to instructions or questions. So I had to figure out how to “program” my chatbot through a bit of trial and error.

    Not all AI/ML algorithms are this complex. Some of them are much easier to understand. It’s a spectrum, and GPT-3 is on the far end of that spectrum.

    Anyway, after a few days of intermittent tinkering, I was able to produce a chatbot that could carry out a sustained and informative conversation about David Hume’s theory of epistemology, to the point where it gave me new insights into the subject. I accomplished this by tinkering over a few days, as a layperson, using plain English.

    I had reached the miracle stage of AI/ML.

    But there were problems. First of all, the chatbot would end the conversation just when it got really interesting. It suddenly would insist on saying goodbye and could not be persuaded to continue talking. It turns out that this kind of AI model has a strict memory limitation. When you hit it, the chatbot suddenly forgets your entire conversation. My new philosophy tutor friend also sometimes gave weird answers. I knew when to ignore them but they would have confused some students.

    GPT-3 has a community of developers who are incredibly helpful, particularly when the topic is something idealistic like education. I was able to find a very knowledgeable programmer who was generous with his time and helped me understand what I would need to do in order to take my tutor to the next level.

    The grind

    The first thing I’d need to do is learn to program in Python since I had reached the limits of programming in plain English. And for the 9,781st time, I was momentarily tempted to learn a little programming. But then he explained what I’d need to do.

    For the memory problem, I’d need to chain portions of the conversation together. But since GPT-3 doesn’t actually remember large chunks of information so much as it distills them, the chatbot wouldn’t literally be able to recall our entire conversation. You can quickly see where this could become problematic. If the student says something like, “When you said earlier that…”, it’s hard to predict how the chatbot would respond.

    And so we enter the grind phase. It could also be called the whack-a-mole phase, since you’re finding a problem, writing a solution, and then looking for unintended consequences elsewhere. Also, since the model isn’t knowably deterministic and the questions students will ask also aren’t knowably deterministic, it’s probably impossible to test all the possible scenarios.

    Which is why you don’t see chatbots that are this open-ended and ambitious. Translating that initial miracle into a reliable response is a daunting if not impossible task. Today’s chatbot designers use UX and context tricks to make the inputs from the students more predictable and they also use less complex algorithms with outputs that they can predict and debug more easily. They tend to reserve the usage of models like GPT-3 for limited and specific applications. And even when they’re careful, producing a chatbot that is rock-solid reliable takes a lot of hard work, including difficult debugging that’s often quite different from debugging traditional software.

    This brings us to the final phase: The wall.

    The wall

    Sooner or later you reach the limit of what your tech can do for you. Predicting that limit in advance takes tremendous skill, often requiring extensive domain knowledge of both the tech itself and the problem it’s being applied to, whether that’s detecting manufacturing defects, discovering new drugs, or tutoring students. There are always nooks and crannies of knowledge and skill that are a poor match for the technology’s capabilities or the data it can access.

    Think about spelling and grammar checkers. They’ve been around since 1961, believe it or not. Even as recently as five or six years ago, Microsoft Word’s spell checker was so bad that I always turned it off. Today, I use Grammarly Pro, which checks spelling, and grammar, and now even makes suggestions on effective sentence structures. I love it. It makes me a better writer.

    But it still makes mistakes in spelling. It makes more mistakes in grammar. And it’s writing style suggestions, while pretty good, make the sentence worse or even change it’s meaning fairly often. The reasons for these limitations are often not obvious to the layperson. For example, Wikipedia notes this eye-opening fact about spell checkers:

    It might seem logical that where spell-checking dictionaries are concerned, “the bigger, the better,” so that correct words are not marked as incorrect. In practice, however, an optimal size for English appears to be around 90,000 entries. If there are more than this, incorrectly spelled words may be skipped because they are mistaken for others. For example, a linguist might determine on the basis of corpus linguistics that the word baht is more frequently a misspelling of bath or bat than a reference to the Thai currency. Hence, it would typically be more useful if a few people who write about Thai currency were slightly inconvenienced than if the spelling errors of the many more people who discuss baths were overlooked.

    Wikipedia

    The tech has a non-obvious fundamental limitation. And in this case, not only is more data not better; more data is worse. So the idea that all AI/ML problems can be fixed with big data is flat-out false. Sometimes better data is better.

    Grammar is significantly more complex than spelling and writing for clarity is significantly more complex than grammar. Each of these functions will likely hit a wall. It might not be a permanent wall, since technology improves over time. But it might be, since sometimes the limitation isn’t the tech but the nature of the problem or the data available in a form that is accessible to the tech.

    Ethical implications

    The rush I felt when I learned something about philosophy from the chatbot I wrote myself is indescribable. I was a philosophy major with a particular interest in anything related to the mind or knowledge. While I don’t have an advanced degree, I certainly knew something about David Hume’s epistemology when I started the dialogue. I was certain I was seeing the future.

    And maybe I was. But it isn’t the near future. When I think about the much more mature technology of the grammar checker, I wouldn’t trust it with weak writers, and certainly not with ESL students. The checker would be more prone to make mistakes and the students would be less likely, on average, to have the confidence and knowledge necessary to know when to ignore the machine. In order for me to change my mind, I’d want to see some quality IRB-approved, peer-reviewed studies showing that grammar checkers help these students rather than harm them.

    We’re in a heady moment with AI/ML. I see a lot of projects rushing headlong into heavy use of the tech, often putting it into production with students without the kind of careful oversight necessary to fulfill the EdTech Hippocratic oath: First, do no harm.