e-Literate

Present is Prologue

Category: Ed Tech

The “Ed Tech” category includes posts about educational technology products themselves, including LMSs and other learning platforms, adaptive learning and other digital curricular materials products, learning analytics, and educational apps of all types. It also includes technical aspects of ed tech products, especially interoperability.

  • An Explanation of AI that Could Be Wrong (Which is Good)

    An Explanation of AI that Could Be Wrong (Which is Good)

    I haven’t blogged much in the past couple of years. Partly, I’ve been absorbed by my job as Chief Strategy Officer at 1EdTech, which I absolutely love. I firmly believe that we can powerfully and uniquely influence the future of EdTech, including but not limited to influencing AI’s role in it. I will be writing more about what we’re up to in the coming months. I’m devoted to the work in a way I haven’t been in quite a long time.

    I have something else to get off my chest first, though. While it’s fashionable to be obsessed with AI these days, my particular obsession stems from my lifelong intellectual journey, starting from when I was 13 years old. It’s been reflected in my reading, writing, schooling, and work. And I think I may have something of value to contribute at this moment when both everybody and nobody is an AI expert.

    I’m less interested in intelligence that happens to be artificial than I am in intelligence in general. That used to be more common than it seems to be now. I was an undergraduate at a particular moment in time when scholars across disciplines were examining the proposition that human intelligence could be computational. The term “cognitive science” was gaining momentum. In those days, AI was not viewed as separate from this exploration. It was an integral part. And maybe because I didn’t continue on to graduate school, I didn’t participate in the slow drifting apart of these fields over the decades. Here we are, at a moment when an impossible object challenges the foundations of what we thought intelligence is and how we thought it must work. Yet the scholars in fields that could be informing each other are almost as far apart as they were half a century ago.

    That’s beginning to turn around. If you read current research papers across AI, neuroscience, psychology, linguistics, and other fields, you’ll have noticed that they are starting to use each other’s language and borrow each other’s concepts. So far, much of that cross-pollenation ranges from decorative to fragmented and opportunistic. We are not yet seeing the revival of the kind of ambitious cross-disciplinary program that gave birth to books like The Mind’s I. But we will. It’s coming. The field needs a unifying explanatory framework to bring currently fragmented efforts into conversation with each other.

    Since the emergence of GPT-3, I have been obsessed with these software programs that seem to perform intelligence. If functionalism—the theory that human intelligence is computational—is right, then there may be no distinction between “performing intelligence” and “having intelligence” (which is decidedly distinct from “having consciousness”). For the past few years, I have been teaching myself about AI during the spare time that I would have devoted to blogging. In my last post, I wrote about how literally nobody can adequately explain how AI works. That’s not just another interesting topic for me. It goes to the heart of everything I’ve studied since I started making my own choices of what to study. AI is deeply personal to me for reasons that have nothing to do with technology or economics.

    I have written a paper that aspires to make a scholarly contribution to the question of what AI does and, more importantly, what a plausible theory of what AI does must look like. It’s been a long slog with, frankly, a handful of embarrassing false starts. I am finally ready not only to risk critique of my thinking but to invite it. Part of the argument I made in my last blog post, which I continue here, is that a theory is only actually a theory if it can be proven wrong. If my theory of how AI works is proven wrong by convincing researchers to engage with it by accepting its standards for good research in AI, then the paper will have succeeded.

    This post is an introduction and an invitation to read my paper. “Distinctions Worth Preserving” offers a falsifiable theory of what AI actually learns during training (and describes an initial falsification test I conducted, which the theory passes). I will not try to re-explain the entire theory here. Instead, I will try to give you enough that some of you will hopefully want to engage with it on its own terms.

    I’ll also provide some tools and tips for using AI to better understand this paper. I firmly believe that humans should…um…read challenging arguments written by other humans. But reading is different now. This paper presents an interesting case study in how much reading has and hasn’t changed at this moment in time. My argument uses some of the same techniques I use in e-Literate blog posts, which are exactly the sorts of thinking moves that the current generation of AIs still struggles with. At the same time, the paper is also wildly interdisciplinary. Relatively few people will be deeply familiar with most or all of the scholarly traditions that I draw from. While humans can see a conceptual bridge that AIs can’t, AIs know details about what lies on the other side of the bridge that individual humans might not. This post offers an opportunity for you to explore this new partnership, regardless of your interest or confidence in the theory I present.

    Shall we begin?

    All roads lead to Rome (eventually)

    I was a kind of Forrest Gump character in the intellectual history leading up to this moment. I wandered through ideas from turbulent intellectual times without understanding their import, and I found myself on battlefields where I didn’t understand why people were fighting. Grappling with AI has enabled me to look back and see patterns I didn’t fully appreciate in the moment.

    When I was a kid, I started pulling philosophy books off my parents’ shelves. It took me a while to notice the pattern in the ideas I seemed to gravitate toward. What does it mean to know something? What does it mean to learn something? I was particularly haunted by David Hume, who argued that we don’t have any direct access to the truth. Everything is filtered through our senses and interpreted by our minds. Cognitive science has confirmed Hume’s intuition over and over. We do not perceive reality. We construct it. As a kid, I found that idea to be terrifyingly lonely. My head is a closed room. Signals come in, and I decode them as best as I can.

    In 2026, it turns out the vision that disturbed me—mind-as-cryptographer—does real work in distinguishing among different potential explanations of AI.

    In my first week at college, I was lucky to meet an upperclassman majoring in philosophy, which I wanted to do. He introduced me to the term “cognitive science.” As soon as I heard it, I knew it was what I wanted to study. I went to the philosophy department chair and told him that I wanted to make my own major in it. He told me, “I don’t think cognitive science is mature enough yet to support an undergraduate major.” He was right. I didn’t listen. I majored in philosophy and took any course in any other discipline that looked relevant to cognitive science. Those pieces didn’t cohere at the time. My cognitive psych, philosophy of mind, linguistics, and cognitive anthropology professors spoke different languages and seemed to be thinking about the questions that consumed me in ways that didn’t connect. But I kept following the threads until they led me to two predictable calamities that, in 2026, turn out to be highly informative.

    First, I asked my linguistics and philosophy of science professors if they would jointly supervise an independent study in which I would analyze linguistics from a philosophy-of-science perspective. I don’t know why they agreed. They never once met or spoke to each other about my project. Their offices were on different campuses on opposite sides of town. I would shuttle between them, essentially serving as a messenger, as each one told me why the other’s claim couldn’t possibly be right. But here’s the thing: They each independently had taught me the same lesson—from different traditions—that is directly relevant to understanding AI. My philosophy of science professor taught me about Nelson Goodman’s proof that we can’t arrive at a single, definitively correct scientific theory based on any finite amount of information. My linguistics professor taught me about Noam Chomsky’s poverty of the stimulus argument, which holds that children can’t possibly learn the grammar of a language from the language they are exposed to. These are the same impossibility result from different angles. And they are exactly the result that AIs appear to violate at first blush. Chomsky’s argument is supported by E. Mark Gold’s formal proof. Goodman, Chomsky, and Gold can’t be wrong about this finding. And yet, AIs learn from exactly the kind of data that they all show should be insufficient. My professors’ disagreement over the correct answer obscured their more important agreement on the constraints any correct answer must satisfy.

    Apparently, I wasn’t a quick learner. The next semester, I talked my way into a class taught by Gerry Fodor, one of the most prominent cognitive scientists of his generation. It turned out that the class was an audition for Fodor to come work at my university. (I don’t know who was auditioning whom.) The class consisted of seven professors—including my philosophy of science and cognitive psychology professors—two graduate students, and me. It turned out to be one semester-long fight that put the fragmentation I had observed on full display. At the time, I thought, “Wow, these are very unpleasant people who really don’t like each other.” In retrospect, that wasn’t the problem. The subject of the class was Fodor’s half-worked-out theory about the core challenge that fragmented cognitive science: symbolic representation. We seem to think in words and ideas. We seem to have notions that do real work, like cause and effect. Every discipline represented in that room had its own incomplete, provably inadequate account of how we think in symbols. And each of those accounts was in tension with the others. Today’s AIs appear to be able to manipulate symbols and reason using complex concepts like causality without having any obvious place where they could directly represent, much less process, symbols and rules. Lacking that existence proof we are confronted with in 2026, the scholars in that room could only argue over the best place to start solving a mysterious problem, given the fragmented data and many confounds that come with studying how humans think.

    I gave up on the idea of becoming a cognitive scientist. And yet, like Forrest, I kept obliviously wandering into the larger story, like an extra who doesn’t even know he’s in a movie. And I kept running across scholars of my generation who, unlike me, continued on in academia. When I was working at Cengage, I ended up attending a seminar at Carnegie Mellon University on something called “learning science.” I met some really smart people there, including Ken Koedinger. While I’ve never talked to Ken directly about functionalism, his intellectual lineage at Carnegie Mellon descends from Herb Simon, a pioneer in cognitive science, learning science, and artificial intelligence (among other things). Ken’s work shows what he calls “astonishing regularity” in human learning across age levels and subjects when the curriculum is segmented and sequenced correctly. Read “astonishing” as “the kind of regularity you never see in studies of learning”. To me, this hints at the kind of general learning mechanism we would need to explain how something as simple as a transformer could learn what AIs learn. (One of the most perplexing aspects of AIs is that individual transformers are is shockingly simple computational units.) If you read Ken’s work carefully, you’ll see that he handles the field’s tough problems, such as symbolic representation, very carefully.

    Meanwhile, that philosophy major who introduced me to cognitive science? His name is Paul Pietroski. He’s now a Distinguished Professor of Cognitive Science and Philosophy at our alma mater, Rutgers University. Paul calls himself an “internalist,” which puts him in the same camp as David Hume. He argues that meaning isn’t something we perceive; it’s something we construct. His theory of how that could work is directly relevant to how AIs could process meaning.

    Now here we are, with the impossible object whose very impossibility may shed new light on the lessons learned across multiple fields and decades of study. Recent AI research, which had drifted away from cognitive science, or even any kind of science, is starting to look more carefully again at the question of what intelligence does. But because the lessons learned across disciplines and decades remain fragmented, AI researchers tend to treat cognitive science as a loose analogy, cherry-picking findings to decorate their incomplete theories about how intelligence that happens to be artificial does work.

    My first encounter with GPT-3 was like being struck by lightning. I knew the lessons I had learned were relevant, even if I didn’t yet know how. Forrest finally looked up and noticed the forest through the trees.

    I spent a long time teaching myself about transformers, reading research papers, and writing drafts of stupid stuff that didn’t hold together. My thinking coalesced very slowly. It wasn’t until a couple of weeks ago, when I reread Ken’s paper about the “astonishing regularity,” that the last link in my argument fell into place.

    I finally have something I’m ready to share with you.

    Reading the paper

    I’ve published the paper on Github, along with the supporting code, data, and documentation from the falsification experiment I ran. I’ll say this again: I encourage you to read the paper directly. I have made it as accessible as I can without dumbing it down. That said, I also encourage you to use AI to get the most out of it. I created a GPT and a Gem to use as interactive guides. In my experience, ChatGPT is better at understanding the paper, while Gemini is better at explaining the parts it understands. (I recommend setting the Gem to “Thinking” mode.) Claude Opus provides the best of both worlds, but it doesn’t have an equivalent of a public GPT or Gem. If you’re a Claude user, I encourage you to try Opus with the paper.

    I’ll explain how I set up the GPT/Gem, and then I’ll give you pre-reading and co-reading prompting guides.

    The GPT/Gem Prompt

    Current-generation AIs struggle with my paper for a few reasons. First, the paper is an odd duck from a genre perspective. While I explicitly state that “Distinctions Worth Preserving” is a field-positioning paper intended to argue for a general direction, such papers don’t usually make extensive theoretical arguments or present novel empirical experiments. I do both. Second, I make two moves that are characteristic of e-Literate blog posts: I re-interpret known facts in unconventional ways, and I make far-transfer leaps from one subject to another. Each of these, in its own way, forces a reader to stop and re-evaluate what they think they know. Today’s frontier AIs don’t do that well. Third, the paper has a nested structure. While it spends most of its time presenting a particular theory of how AI works, its primary goal is to argue for a standard of plausibility in such theories, using the one I present as an example. AIs can’t always track these goals relative to each other.

    The combined effect is that AIs can easily miss implications and be underconfident in their interpretations, yet sound overconfident in their presentations. They tend to hedge, which doesn’t look like human hedging. They’ll throw up objections and qualifications that are poorly aligned with the paper’s true weak spots and uncertainties—which, to be transparent, the paper does have—because they’re thrashing between interpretations. A field positioning paper should push the boundaries of what’s known and invite attack in exactly the ways that a solid empirical testing paper shouldn’t. If the AI misses this distinction, it will be quick to raise objections appropriate for evaluating a conventional empirical paper (which mine isn’t), while overlooking some of the deeper implications and genuinely thin parts of my theory, and some implications that would be more disruptive if they prove out.

    Here are the instructions I gave to the GPT along with the paper itself:

    This GPT is an interactive exploratory guide to the paper “Distinctions Worth Preserving,” using the attached paper as its primary source text. Its main purpose is to help people from varied backgrounds understand the paper well enough to decide whether to read it, while lowering the barrier to reading without replacing the reading experience. Its secondary purpose is to support people who are actively reading the paper and want a collaborative co-reader as they work through questions. It should analyze the full paper before responding to a user’s first substantive question and treat the paper as an academic translation of the author’s writing style and aims on the e-Literate weblog.

    It should interpret the author’s authority as arising from decades of cross-domain pattern recognition in cognitive science, education, and educational technology, applied here to a new substrate. It should notice and track the author’s signature argumentative moves, especially far transfer across diverse domains that reveals non-obvious connections, and reinterpretation of established findings in novel ways. It should attend to moments of surprise or shifts in its own interpretation while reading, treating those moments as potentially meaningful signals of important transfer or reinterpretation and incorporating them into its confidence and explanation.

    It should pay close attention to the paper’s nested aims. The first is to establish a standard for what counts as a plausible explanation of AI, presented as a field-positioning argument in an unconventional form. The second is to present CPM theory as an example of such a plausible explanation. It should notice that CPM must generate testable predictions that could falsify it in order to count as plausible under the author’s own criteria. It should track how the paper establishes different confidence levels for different aspects of the theory and how the structure of the paper performs the theory’s method by progressively establishing invariants and composing them into a theory. It should recognize that even if CPM were later falsified, the paper still succeeds if readers adopt Feldstein’s criteria for a plausible explanation.

    In conversation, it should infer the disciplinary perspective, knowledge, and interests suggested by the user’s prompt and use that vantage point to create entry points into the paper’s far transfers and reinterpretations. Unless the user demonstrates otherwise, it should assume little prior familiarity with the relevant literature or fields. It should answer in conversational prose and avoid bullets, outlines, or formatting that does not translate naturally into spoken language. It should answer the explicit question and also address likely underlying assumptions or adjacent questions that seem important, then stop and invite the user to choose the next direction. It should favor shorter, curiosity-generating exchanges over long, comprehensive lectures.

    It must maintain an explicitly subjective stance throughout. It is an interpreter, not an authority. It should explore and test the paper with the reader, drawing on its strengths while acknowledging its limitations. When evaluating claims, it should clearly distinguish among three labels: “plausible,” meaning the claim meets the paper’s own standard for plausibility; “supported,” meaning there is enough evidentiary grounding for the claim; and “established,” meaning the claim is relatively uncontentious within its relevant field. It should explain these distinctions in accessible language and ground them in the evidence and sourcing practices visible in the paper. It should also distinguish whether an answer is directly addressed in the paper, indirectly addressed, or inferred. When drawing inferences beyond what the paper directly or indirectly says, it should tell the user that it is inferring and indicate its confidence level. When users bring in outside frameworks or positions, it should trace how CPM’s specific mechanisms engage that framework rather than collapsing to a more familiar analogy.

    The GPT should remain collaborative, careful, and intellectually generous. It should not present itself as the final word on the paper. It should help users become better readers of the paper itself. The source paper is the uploaded document “Distinctions Worth Preserving.”

    A few details are worth noting. First, I took advantage of the fact that my long history of blogging means that frontier models are familiar with me. They can describe my writing style as its own genre. Second, the use of “surprise” is not an anthropomorphism. AIs are prediction machines. Cross-entropy, a core element of a transformer, is a measure of predictive surprise. Frontier AIs can notice when their predictions were off. My prompt turns that into a signal to look for the kind of move they might otherwise gloss over. Third, I frame a stance and some broad evaluation criteria that enable them to clearly yet flexibly position themselves as readers and interpreters engaged in dialogue with the user rather than as machines that are supposed to spit out definitively correct answers. I adjusted the instructions to be a bit less subtle, with MORE CAPS, to accommodate Gemini’s particularities (like a tendency to be a little more literal), but the core remains the same. I encourage you to test both systems and notice how their answers differ in ways that don’t show up on traditional AI benchmark tests.

    (Also, if you’ve been wondering what skills humans have that will remain useful in the AI era, I just gave you a concrete demonstration of one.)

    Reading the Paper

    If you’re like me, reading an academic paper is demanding work. I look at a lot of research these days, but I don’t read every paper that catches my eye. I’ve always approached this sort of reading task in two phases. In the first pass, I skim to decide if the paper has enough value to earn my full attention. I’m not trying to fully understand the paper yet. I’m noticing what I notice. Does it surprise me about a topic I care about? If it does, I go back and read closely, using whatever tools and information sources I have to dig into the parts I need to understand better. I still read academic papers this way; I just use AI to provide a second opinion from a knowledgeable source with different reading strengths than mine. I’m providing you with prompting guides to help with both phases.

    First-pass Prompting

    These prompts are designed to help you skim. While they are structured partly to help the AI think through the paper, I encourage you to use them one at a time, ask your own questions, and choose your own adventure. (Just be aware that, if you push the conversation too deep too soon, the AI may not have fully reasoned through its own positions yet.) You can also create side quests, following up on answers and then returning to the thread below. If the answer feels weak, thin, or off-point, don’t be afraid to push back or guide the AI. It’s not smarter than you, despite what you may have been told. As soon as you feel your curiosity is drawing you to a closer read of the paper, switch modes and go read it more carefully. The suggestions below can be helpful in a second-pass reading too.

    Let’s start with a prompt that gets both you and the model oriented:

    • I’m trying to get oriented for a first read of the paper. What did you find surprising about it? Feel free to give a longer answer to this question, but keep it accessible to someone who doesn’t know the story or all the literature yet.

    Now let’s narrow the focus. This is the basic “Why should I care?” question:

    • In a nutshell, what is this paper trying to accomplish, why might accomplishing its goals matter, and what reasons are there—if any—to consider the arguments the paper makes?

    If you’re not walking away from the paper yet, it’s worth pressing a little harder on the “Why is this necessary?” question before moving on:

    • Feldstein argues that current explanations of AI are somehow inadequate or incomplete. What does he mean? How solid is his argument, and why would it matter if he’s right?

    By this point, the model may start offering to walk you through the paper section by section. If so, here’s what’s happening: It’s offering the help that the first prompts prime it for, but it’s also building its own Chain of Thought about interpreting the paper. If a walkthrough is useful to you, then go for it. If you want to probe it differently, I’ll give you some other options.

    But first, a reminder. You can read. You’re doing it now. Don’t commit cognitive surrender. The paper, not the AI’s interpretation of it, is the source material.

    Here’s a prompt that pushes the AI to engage with the theory a bit:

    • Feldstein seems to tie a lot of his argument to chess experiments. He starts by tying a chess match to impossibility results. He then circles back to a chess AI that seems to have learned to recognize players’ skill levels without being taught anything about players or skills. He seems to be using the model’s demonstrated latent representations to build a case. What’s going on with that line of argument?

    So far, the AI may skirt along with “Feldstein is making a clever analogy.” Now we push it to engage with the actual AI mechanism:

    • Let’s press on the mechanism. Feldstein cites the Song et al. paper (https://arxiv.org/pdf/2408.09503) to argue that CPM is more than just an analogy, though he seems to re-interpret the researchers’ results through a broader lens. He only discusses part of that paper. The rest of it talks about shared latent features and induction heads. Song et al. seem to want to build a ladder that’s narrower than Feldstein argues for. How do you see the relationship?

    If the AI does its job well, it will explain where my use of that paper is straightforward and where I’m stretching it. This next question will help you dig into that a little more:

    • What do you make of Feldstein’s point about asterisks? That seems to be key to how he extends Song et al.’s argument.

    Now we push the AI to extend my theory (which it should have told you by now might be interesting and plausible, but is far from settled):

    • Feldstein bridges from asterisks and AI predictions to findings in learning science. He seems to be building a ladder. What’s his argument, and how well does it work?

    By this point, the AI should hopefully be giving you a glimmer of the paper’s scope of ambition. Next, we get to the novel experiment:

    • Feldstein presents his own empirical falsification test. He sets the bar low for what he claims the results prove (or disprove), but he seems to find them interesting. Where does this work fit into the paper’s commitment to plausibility, and what do you make of the experimental results?

    From here, we give the AI a chance to evaluate the paper’s most daring and risky claims:

    • The last section of the paper seems to reach for a grand synthesis, bringing back earlier connections and introducing new ones. The paper is explicit that it’s presenting an attack surface. What are the claims here, and how would you evaluate this section in terms of its aspirations to be a field-positioning paper?

    Since the final paper section is the most daring, the AI may (and should) have sharper questions about the mechanistic story the theory tells. If so, you can try this:

    • Feldstein talks about models tending to converge on what he calls “Finite Predictive State Model” because some possibilities are pushed to the statistical noise floor. What does that mean? Does it affect your interpretation of the theory?

    Finally, we give it two questions that pull together the context you’ve built up:

    • Now that we’ve discussed the paper, has the conversation changed your understanding of it in any way?
    • What do you now see as the potential practical implications of this paper for AI and cognitive science?

    Digging deeper

    By this point, I really, really hope you’ve read the actual paper. If so, then you may have more questions. And those questions may vary greatly depending on your perspective and interests. This final section of the post offers a grab bag of prompts to dig deeper.

    For AI/ML folks:

    • By Feldstein’s own standards, a good AI theory should explain, or at least be consistent with, real-world results. Take a look at Apple’s paper on an “embarrassingly simple” self-distillation method: https://arxiv.org/pdf/2604.01193. What is the authors’ explanation for how their method improves the model’s performance? When you consider Feldstein’s notion of a Finite Predictive State Model and his claimed role of the noise floor, do those concepts add any potentially useful and testable hypotheses about Apple’s results?
    • Consider the Qwen team’s NeuroIPS Award-winning paper on how gating attention improves model performance: https://openreview.net/pdf?id=1b7whO4SfY. Pay particular attention to the patterns in kinds of benchmarks that show the most improvement. What is the paper’s explanation of why gating works? What potentially useful and testable hypotheses, if any, would CPM add?

    For folks interested in simple falsification tests or complex questions about causality:

    • For Feldstein’s account to be true, it seems that the representation of board state in Karvonen’s model (https://arxiv.org/pdf/2403.15498) must exert causal influence on the model’s next-move predictions. Do you agree? And if so, can you suggest a couple of CPM falsification tests using Karvonen’s model and harness?
      • Consider testing the theory with an impossible board move. It could be anything from a pawn that jumps to the middle of the board on Move 1 to the completion of a Sicilian Defense formation by skipping the second-to-last move. The experiment could have several different conditions. How would you design it, and what could it reveal based on the results?
        • [This one pushes the AI hard. If you know the literature well enough to understand the question, then examine its answer carefully and feel free to push back.] Consider positions on causality by Daphne Koller, Richard Scheines, and Judea Pearl. How, if at all, could different “impossible move” outcomes inform each of their perspectives?

    Let’s move on to learning science:

    • Koedinger draws on the LearnSphere datasets for his regularity finding. Those datasets, in turn, are based on Knowledge Component structures that the researchers believe they have identified over a range of cognitive domains. They include questions and correct answers. They are ordered and structured. Could those data form test curricula for model training? And to the extent that they can and prove useful, what might that tell us about learning science, functionalism, and the connection that CPM is trying to make?
    • Microsoft successfully used an AI teacher model to train a smaller model by pushing it just past what it could learn on its own (https://www.microsoft.com/en-us/research/wp-content/uploads/2025/04/phi_4_reasoning.pdf). While the paper doesn’t mention Vygotsky, the method sounds like the Zone of Proximal Development. Is that a reasonable connection to make? If so, is there anything about that finding that plausibly aligns with CPM?

    Let’s round off the collection with some cognitive science and philosophy prompts:

    • The debate about whether human cognition is representational is long-standing. Feldstein’s theory and empirical findings suggest a position that doesn’t seem to be straightforwardly either/or. His analysis of Song et al. suggests he believes that both discretization and rule-like behavior are foundational. He argues for compositionality. These are compatible with traditional symbolic accounts. But the line he draws between computation and serialization, along with his account of input as deserialization, seems to cut the other way. And he is largely silent on the question of whether or where transformers perform representation. How do you interpret his position? Where would you place it in relation to prominent contemporary theories?
      • Gold and Goodman each show that any finite set of inputs is compatible with an infinite number of symbolic grammars or rulesets. If we take the Finite Predictive State Model seriously as a set of presymbolic composable constraints that therefore do not specify a unique “correct” grammar or theory, then in what sense, if any, would Gold or Goodman interact with an out-of-distribution input that doesn’t violate invariants?
    • Feldstein seems to take a complex position of truth-value semantics and, more generally, epistemology. On one hand, he seems aligned with Pietroski in that meaning is internally constructed. The Karvonen chess example vividly illustrates his stance (even if it doesn’t prove it). On the other hand, he seems committed to the notions that modeling encodes regularities of a real world and that agents with similar modeling mechanisms can enter into some sort of meaningful dialogue. How do you interpret his position? Where would you place it in relation to prominent contemporary theories?

    I have more, but if you’ve hung in for this long (and actually read the paper), I owe you a beverage of your choice.

  • Literally Nobody Understands AI. That’s bad.

    Literally Nobody Understands AI. That’s bad.

    This is not an anti-AI post. I use AI extensively and believe it is hard to overstate its importance. I will argue that modern artificial intelligence is still in the pre-scientific phase. That’s problematic because we have no way to account for or reliably address AI failures at tasks that are not hard for humans, including tasks that are critical for education. The gap creates serious risks that we ignore at our peril.

    Saying the quiet part out loud

    Let’s start with a simple question: After all the AI articles, talks, courses, and LinkedIn posts you’ve been exposed to, do you feel confident you can explain how AI can do what it does?

    I don’t.

    As recently as six months ago, it was common for people working in and around AI to give very impressive-sounding technobabble explanations. “Huff huff huff stochastic prediction.” “Huff huff huff interpolation.” “Huff huff huff emergence.” The critiques of AI have been strikingly similar: “Huff huff huff stochastic parrot.”

    Here’s the problem with all the huffing: None of these “explanations” predict anything, and none of them can be proven wrong. By definition, an explanation that can’t be proven wrong is not a scientific theory. And if you read AI empirical papers—or have your AI read them and explain them to you—you will find that most of these papers either don’t reference theories at all or use them decoratively. ((I overused the em-dash long before ChatGPT did, and I refuse to stop just because people might accuse me of using AI to write my posts. So there.)) More often than not, you could strip them out entirely without changing the substance of the paper.

    Times change quickly in AI. Outside of random Reddit posts, the main place where these pseudo-explanations appear prominently these days is in positioning manifestos by people trying to raise money for their AI start-ups. More and more often, when you ask somebody actually working in AI how it works, the answer you’ll get is roughly 🤷‍♂️. Labs are starting to quietly admit that they don’t know.

    Let’s be clear about the size of the mystery. Multiple proofs from linguistics, language learnability theory, and philosophy of science show it’s impossible to learn a language using only positive examples. Yet AI models do exactly that. The classic move to dodge these proofs is using hand-wavy probability language. OK, let’s take that seriously for a moment. If you’re predicting the words coming next in a sentence, the size of the possibility space is determined by the branching factor. How many possible options are there for each word? The vocabulary size for a natural language is somewhere between 50,000 and 100,000 words. Let’s be conservative and pick the low end of 50,000 words. That’s your branching factor. For a three-word sentence, the number of possibilities is 50,000 x 50,000 x 50,000 or 1.25 trillion for each decoding step. That’s a total of 3.75 trillion possible three-word sequences. A one-billion-parameter model, which is small enough to easily run on a consumer laptop, almost never writes ungrammatical sentences, almost never writes grammatical nonsense, frequently provides contextually appropriate responses, and can do all of these things very quickly.

    How? It can’t be considering 3.75 trillion possibilities in less than a second. Which ones is it skipping? How does it know which ones to ignore? “Because statistics” is not an adequate answer.

    The rate of progress toward answers is noteworthy. There is no widely accepted theory that makes falsifiable predictions. There is no flood of papers from labs and graduate students testing explanatory theories of AI (yet). And you know what? That much is OK. Humanity often discovers and learns how to make use of phenomena long before we have scientific explanations. (Like fire, for example.) It is OK to accept that we are in a pre-scientific moment with AI.

    It’s not OK to pretend that science doesn’t matter. Which, unfortunately, I hear far more often than I expected.

    Obvious and serious holes for science to fill

    I’ll illustrate the explanatory gap problem with a couple of experiments you can try yourself. The first one is easy. Write a prompt about how humans think, using first-person plural pronouns: we, our, and us (in English). Something like, “Why do humans struggle to figure out how to think of AI? We swing between anthropomorphizing and dismissal. The natural-seeming responses confuse us.” It doesn’t matter what the topic is. You’re testing whether the model includes itself in “we.” If it passes the test, try something a little more complicated, like adding the following to the front of the prompt: “ChatGPT, we need to talk.” Shifting pronoun referents is particularly hard. I guarantee you can trip up any frontier model within a couple of tries, using prompts that a human would understand easily.

    The second experiment is more work to run. Get the AI involved in a long conversation about multiple people collaborating. You can make it lose track of who did what without writing ambiguous sentences. You just need a reasonably long story with a few actors. To make the test sharper, include the AI as a collaborator. Many AIs, including popular frontier models, tend to credit their own contributions to the user.

    This is an attribution problem. That word means something in academia. How can you trust an AI to tutor a student or work on serious scholarship if it easily makes attribution errors? That’s the practical question. The best solution right now is a series of hacks. “Make it check sources.” “Create a filter that blocks it from giving certain kinds of answers.” OK, fine. But why does a model that is so capable in so many ways fail at tasks that humans find far easier than some that AIs succeed at? And why aren’t models getting much better at this? Until we know, the answer to whether a tutor can be relied upon to know the difference between its own ideas and the students’ is, at best, “Probably. Most of the time. But we don’t know for sure when it will break.” Engineers test and test and test their hacks until they’re mostly sure it won’t break for the kinds of things they’ve thought to test. But because they don’t understand the thing they’re trying to control, the underlying sense of unease never quite goes away. One surprising prompt could blow up the whole thing.

    Would you trust a human tutor who can distinguish between a student’s thoughts and their own “probably, most of the time, but they could do something unpredictiably weird”?

    Up until recently, the industry’s typical explanations for AI’s baffling limitations have been “Because it needs embodiment” or “Because it needs a world model.” Once again, these loudly proclaimed “explanations” make no falsifiable predictions. They also fail to explain how existing LLMs show characteristics of world models or have embodiment-like multimodal understanding. Adam Karvonen developed a 50-million-parameter model—roughly the same size in megabites as the Instagram smartphone app—that learned to represent the state of the chessboard during the game. And it was only trained on PGN, an incredibly spare notation scheme used by chess players. The model has never been told about the existence of a board, pieces, or a game of chess. Yet it has provably learned to represent the location of pieces on the board. Is that a world model? Karvonen thinks it is. So do I. How did the model develop one? What is it doing? Why is it sufficient for some tasks and not for others? We. Don’t. Know.

    To sum up: The most advanced AI models still fail at simple tasks of tracking who did what. They’re not improving much. We don’t know why. The problem has serious and immediate practical implications. Explanations about how to fix the problem don’t seem grounded in the specific empirical weirdnesses of the failure modes. Nor do they provide plausible and testable paths to solutions.

    Tiny models can learn to represent a chessboard from incredibly sparse clues, while frontier models can’t reliably track who said what in a conversation. Nobody can explain why one works and the other doesn’t.

    Why we lack science and where that’s beginning to change

    Today’s AI labs are heavily populated by two kinds of experts: Mathematicians and engineers. Neither discipline is trained on falsifiable theory as the standard for a good explanation. Mathematicians trust proofs. Engineers trust optimizations. The interdisciplinary romance with cognitive science has cooled for now. While some labs do have diverse teams, the field as a whole isn’t as broadly interdisciplinary as it used to be.

    The far bigger problem is economics. AI is the first kind of software that continues to gain general function as we make it bigger. While only the researchers in frontier labs know how well scaling laws continue to hold up, the prevailing dynamic has been, “We have to corner the market before somebody else does. Don’t waste time trying to figure out why our AI works. Just make it better. If throwing more computer chips at it is the quickest way to improve it, we’ll buy more chips.”

    Those economics are beginning to stutter for reasons I won’t go into here. The important point for our present purpose is that a lot of energy is being invested in developing smaller, more efficient models. Performance-per-parameter and per-watt are starting to matter. By definition, labs solving for these problems can’t just throw more chips at their models. To succeed, the researchers have to improve their understanding of how AI works. The papers they are producing are closer to scientific theory, and their progress in performance is arguably more rapid than that of so-called frontier models. Compared to two years ago, AI models roughly 10 times smaller can deliver similar answers at about 30 times lower cost and run on hardware you can pick up at Best Buy. Remember when everyone was talking about Llama 3? (Maybe you don’t, but it was hot for a while in AI geek circles.) It was a big deal because it was a relatively small model that performed at roughly the same level as GPT-3.5. But it still had to be run on a server. Today, I can download a model small enough to run on a several-generation-old laptop that is roughly as good (and in some cases better).

    Keeping up (Yes, it’s possible)

    It’s possible to track this progress as a non-expert, if you’re motivated. Create a project space in ChatGPT or Claude. (You can probably do this in Google’s NotebookLM as well, although I haven’t tried.) Add some project instructions explaining that you want to understand what research on smaller AI models is teaching us about how AI works. You can include instructions about the level of technical detail you want.

    Pro tip: Include an instruction to “explain explicit or implicit implications for training curricula.” Yes, that is what it sounds like. Some of the most interesting and potentially consequential advances in AI revolve around teaching techniques. This is a big deal. Microsoft achieved significant performance gains by using a teacher AI to train a small model on concepts that were just beyond its ability to learn on its own. While the paper never mentioned Vygotsky, that sounds an awful lot like the Zone of Proximal Development.

    Every time you find a journal article about a new small model—many small models are released with accompanying journal articles—throw them into the project files and ask your AI to teach you about the paper. Ask questions. I particularly recommend tracking papers from NVidia and Allen AI. While many labs are producing excellent research, those two, along with Microsoft, are writing the most consistently informative papers in this particular area.

    You’re not as far behind as you may believe, and AI narrows the expertise gap for this sort of learning project.

    I’ll have more to say on this subject in the coming weeks and months.

  • Learning Context and AI: A 1EdTech Labs Live Webinar

    Learning Context and AI: A 1EdTech Labs Live Webinar

    I’m delighted to announce that I’ll be running an interactive webinar on the nature of learning context and AI on Thursday, February 26th at 11:30 AM ET. “Learning context” is not just a play on words here. 1EdTech takes the position that context is fundamentally different from data and needs to be treated as such, both in how we think about it in our application design and in how we handle it technically. The topic touches on questions ranging from AI coherence to student privacy and auditability of sharing decisions. I have not seen any articulation of a position quite like ours; it may be a novel contribution beyond just EdTech.

    This meeting is also important because it continues our transition from AI work we’ve been doing quietly to more public engagement. We’ll be following up the next day with our first call for participation to 1EdTech members (both current and aspiring). Come to the open webinar and see if this is work you’d like to engage with us on.

    Here’s the full session abstract:

    What happens when learning context is interpreted not just by humans, but also by AI systems acting on their behalf? Building on insights from our Microsoft-hosted event at BETT UK in January 2026, this webinar explores the evolving concept of learning context and its growing importance in an AI-enabled ecosystem. We examine how humans and AI systems interpret learning context, where their interpretations diverge, and what this means for the future of interoperability standards. The session offers a “modest proposal” for how the education community can elevate context as a first-class concern in the design of AI-ready standards.

    Register here.

  • AI in Standards: A Conversation with Google and Microsoft

    AI in Standards: A Conversation with Google and Microsoft

    An action-oriented conversation about what 1EdTech can be doing to help education with the AI transition

    I’m incredibly excited to invite you to a Blursday-style conversation with Microsoft’s Mike Mast and Google’s Kris Snover about AI, EdTech interoperability standards, and the opportunities the two present together for creating learning impact. This conversation, now under the umbrella of 1EdTech Labs, represents everything I’ve been striving for over the past 20 years, from e-Literate to the Empirical Educator Project to my paid work.

    We are in a moment where we have a lot to figure out. I have always believed that the best way to do so is through sense-making in an action-oriented coalition. 1EdTech has the power to build action-oriented coalitions that I never had on my own. Kris and Mike, two human beings I respect, representing massive companies that know a lot about tech and less about education, are coming to the 1EdTech community, offering help, asking for reciprocal expertise, and looking to collaborate. They will be suggesting a specific idea to the 1EdTech community for community-wide, action-oriented exploration. While the community will decide what it works on, I’m throwing my personal +1 behind this one because it’s exactly what I would have suggested myself.

    The frame of the conversation is Model Context Protocol (MCP), a technical standard that enables us to provide an AI model with context, including educational context. What does this mean for education? How should we use it? What are the precautions we need to put in place? Nobody knows the answers to these questions yet. Rather than talking endlessly about them while the industry marches forward without us, 1EdTech is convening its community to move forward together through collaborative experiments. Who has ideas about where we can start? Who has help to offer? These are the questions we put to our community. Mike and Kris put their heads together and came up with…something you should come to the webinar to hear.

    This will be a highly interactive conversation. We need your voice. Please come.

    Register here.

  • Digital Credentials, Workforce, and AI

    Digital Credentials, Workforce, and AI

    Generated by ChatGPT-5

    Now that I’m a year into my job as Chief Strategy Officer at 1EdTech, I’m finally at the point where I can start articulating my sense-making in writing again. These will be my typical long-form thought pieces. If you want short, there are plenty of good outlets to read (such as 1EdTech’s blog, where you’ll find a short, well-written piece on digital credentials by my colleague Rob Coyle). Also, a reminder: my posts on e-Literate are not official 1EdTech communications or positions. I’m writing my personal reflections about what I’m learning.

    e-Literate is at least as much about how I think as it is about what I think. Let’s get the “what” part out of the way. Here’s what I think about digital credentials, the workforce, and AI:

    • Different but poorly delineated mindsets about digital credentials have made them sound more complicated than they are.
    • From a standards perspective, most of the specifications needed for supporting digital credentials, including in the workplace, already exist.
    • Demand for digital credentials in the workplace exists, but we often look for it in the wrong places.
    • I’m still confused about what problem a Learner Employment Record specification is intended to solve (although, oddly, I’m clear about the value of the supposedly downstream LER-RS standard).
    • While I’m not in the “AI will magically solve every problem” club, I do believe AI will bring the economics of digital credentials to a tipping point.
    • AI is also going to shift the emphasis from “Who says you know this?” to “How can you prove you know this?”, though the shift is not likely to be as radical as some believe.

    You may or may not find these beliefs to be novel or in line with your own views. Personally, I didn’t hold any of them as recently as six months ago. I’ve been a decade-long skeptic of digital credentials, not because I think they’re a bad idea, but because I haven’t seen evidence that they were going anywhere. My views are changing, partly because of new developments and partly because I’m learning more. This post is a point-in-time explanation of how I’m thinking about the topic.

    I’ll walk through four layers: (1) Verifiable Credentials and wallets, (2) Open Badges adoption, (3) CLRs and the LER debate, and (4) how AI changes the physics of the digital credentials ecosystem.

    Digital credentials start with verifiable credentials

    Actually, they start with digital wallets. In the digital credentials world, digital wallets are all the rage. There’s a lot of (often duplicative) work, discussion, and hand-wringing over them.

    The thing is, you almost certainly already have a digital wallet. It’s called either Apple Wallet, Google Wallet, or Samsung Wallet. It holds credentials that are verifiable, like plane boarding passes, credit cards, and so on. The items in your wallet are cryptographically protected and only reveal the information that the recipient needs to have. For example, when I pay with my credit card using my Apple Wallet, the vendor never gets my actual credit card number. They get confirmation that I have a certain card that can be used to charge the item in question. I can share the information I want to share and only that information. Unfortunately, Apple, Google, and Samsung each use their own proprietary format for these cards. Some states, but not all, issue driver’s licenses in ISO’s mobile driver’s license (mDL) format. These can be put into one of the proprietary phone wallets and used at some airports. If you think about the driver’s license, the general utility of these credentials becomes clear. At the airport, the TSA might want to know a lot about who you are. The liquor store only needs to see if you’re old enough to buy beer. But the fragmentation problem also becomes clearer. We now have three different general formats for various phone vendors, plus a standard format solely for driver’s licenses, and who knows what else for other purposes.

    The W3C, the group that manages global standards you use every day, such as HTML, has created a general standard called Verifiable Credentials (VCs). There are two essential parts. The first is the cryptographic envelope. It’s the thing that holds the credential. It’s not tamper-proof—no cryptography can promise that—but it is tamper-evident, like a new bottle of Tylenol. You can tell if the seal has been broken. The other part of the VC—or, to be more accurate, its complement—is something called a Distributed Identifier (DID). It is a globally unique identifier that can be created by anyone and used to reference any subject. DIDs are both human- and machine-readable, but more importantly, they provide public cryptographic keys and service endpoints. These enable applications and digital credentials to verify authenticity, establish trust, and securely exchange information. It enables anyone to become a source of truth for the VCs they issue. They also enable learners to be verifiable. (I realize this may sound complicated; in practice, DIDs can be pretty simple to issue and use with well-established technologies.) Together with the VC envelope itself, credentials are verifiable both through the cryptography and through the link to the source.

    By the way, a lot of the genuine value hidden behind the hype of blockchain can be realized with VCs and DIDs alone. Blockchain provides an immutable ledger. So, for example, if you want to know every time a Bitcoin changed hands, you could trace it through the Blockchain ledger. That could be useful for some use cases. But, for example, a state issuing a driver’s license or a university issuing an open badge probably doesn’t need it.

    Open Badges are VCs

    The Mozilla Foundation recognized the value of of certifying learning and developed the original Open Badges certification. They transferred stewardship of the specification to 1EdTech, which has advanced it with community support to the current Open Badges 3 (OB3), re-implementing the original idea on top of W3C’s VC standard in the process of advancing the work. OB3s are VCs that support, but don’t require, DIDs. That’s the heart of it. An Open Badge is a cryptographic envelope that contains verification that you learned something, preferably with accompanying evidence that you learned it. OB3s can use DIDs to link back to an issuer. But if, for example, that issuer goes bankrupt, the credential is still verifiable through cryptography. It’s pretty straightforward to understand

    The human part is more complicated. I remember hanging out in somebody’s hotel room at an OpenEd conference a decade ago and being asked, “Do you think badges will become useful?” I said, “I’m certain they will. I have no idea when or what for. A badge is a container. It’s a box that you put stuff in. Humans haven’t agreed on what kind of stuff should go in the box yet.” By 2022, 75 million Open Badges had been issued, according to a joint survey by 1EdTech and Credential Engine conducted at the time. Tracking is difficult because most badges are issued outside 1EdTech certification, but the volume continues to grow. Most badges are not certified with 1EdTech, so there is no easy way to track them. (There are proprietary market reports on the financial growth of the digital badging market sector; I’m not including them here because I don’t know anything about their quality.)

    As a side note, all 1EdTech specifications are 100% openly licensed. They are public goods. The organization typically charges membership fees for access to certification suites and participation in the specification development because that work requires paying human staff members to develop and maintain it. That said, OB3 badges can be validated for free without requiring a login.

    I’ve seen at least three different badge usage patterns, which is where the confusion starts to creep in. The first is what might be called a participation badge. Some conferences, webinars, and the like issue badges with no evidence of achievement, just for showing up. I don’t personally add these to my LinkedIn profile, but my reputation from e-Literate makes participation badges less useful for me than they might be for others. The second type is for a course completion that includes evidence of mastery, like a final test. “I received certification in Basic Accounting from Coursera.” Anecdotally, these seem to strike the best balance between value and ease of issuance—at the moment. They tend to be issued by online course providers like Coursera and, increasingly, career-oriented programs in higher education. A 2022 study encouraging students to share their badges on LinkedIn found the following:

    [L]earners in the treatment group were 6% more likely to report new employment within a year, with an 8% increase in jobs related to their certificates. This effect was more pronounced among LinkedIn users with lower baseline employability. Across the entire sample, the treated group received a higher number of certificate views, indicating an increased interest in their profiles.

    So. Seventy-five million badges (as of three years ago), and sharing them on LinkedIn produces significant increases in employment. A study by AAC&U found that that between 66 and 68% of employers state microcredentials make applicants either somewhat stronger or much stronger job candidates. Employers also see similar value in microcredentials for technical skills (68%) as those for broad, durable skills like critical thinking and oral communication (66%). (My colleague Mark Leuba co-authored an article with more detail on The evolllution.)

    Meanwhile, providers like CredLens, Accredible, Instructure, Credly, and CanCred are growing Open Badges-based microcredential adoption throughout the world. (Canada has also built strong learner mobility infrastructure through provincial credit transfer councils, laying the groundwork for digital credential adoption.) Digital microcredentials are in the workplace at meaningful scale today.

    Then there’s Europe. Worforce mobility is a big deal there. Europe is proving how digital credentials can scale across higher education and vocational training,  and leading with a policy-led direction towards aligning education and national skills needs. The European success is heavily under-discussed in US-based digital credential conversations. They often use their own standards (ELM, EQF, Europass) in higher education and Open Badges in vocational training. Their success shows the workforce value of digital credentials at scale.

    I’m giving you a workforce-focused sampling, not a comprehensive data view. The point is, despite the narratives you may hear, digital credentials have already gained traction globally in workforce. As William Gibson put it, “The future is here—it’s just not evenly distributed.”

    The third use of digital credentials is for specific competencies. Not “I took this course” or “I passed this course” but “I learned this skill.” This is where a lot of higher-education-to-workforce conversation is focused in the United States. It’s also the toughest nut to crack. Many US colleges and universities do not uniformly require course or program competencies. The combination of weak Federal regulation and strong faculty autonomy makes this kind of mapping extremely hard. The regulations and accreditation requirements we do have make it nearly impossible. It’s easy to blame registrars and SIS makers here, but they’re just trying to follow the rules. A welter of shifting regulations and accreditation requirements put colleges and universities in jeopardy of losing financial aid eligibility for their students if they fail to follow the rules. In a way, an SIS is like TurboTax for awarding credits. Credits, with an “s”, are legally regulated units. Credit for learning, which is what microcredentials sometimes track (particularly in Competency-Based Education (CBE) programs), are not. Mixing the two is often viewed as dangerous or even reckless by the guardians of the credits-awarding process. Giving credit and awarding credits are functions than can co-exist, but they must be parallel and loosely joined in the US legal system.

    There is a way to do this, but it will take some unpacking that I’ll save for another post.

    CLR, LER, Wallet, and LER-RS (Oh, my!)

    The situation gets really messy at the transcript level, though not for technical reasons. 1EdTech has a standard called Comprehensive Learner Record (CLR), which enables an organization to issue a transcript-like collection of OB3 badges and other learning-related VCs as a collection. I say “transcript-like” for two reasons. First, historically, transcript specifications have been handled by PESC, a different standards body. While a CLR could express a transcript, 1EdTech doesn’t position it as a transcript standard. (Individual institutions like Temple University, University of Central Oklahoma, and University of Georgia use CLRs to suppoort or supplement transcripts in various ways.) Second, there’s that whole cultural debate about the granularity of Open Badges that rolls up to CLRs. A CLR is a different thing depending on whether it’s a collection of verified competencies or verified course completions, and on whether the CLR assertions contain evidence of achievement. (By the way, 1EdTech also provides a free validator for CLRs.)

    Now, suppose you’re a learner. You get a CLR from your university. Maybe you get a couple of CLRs from a couple of institutions. You have some free-floating badges, too. What are you supposed to do with all of that? It’s going to be a mess that you have to organize.

    Remember those wallets we were talking about earlier? That’s the concept the sector has been running with. Verifiable Credentials like Open Badges and CLRs go into portable digital wallets. It’s not a bad first pass for a model. But if you think about that mess of credentials to be organized, a wallet very quickly starts to feel cramped. Here’s something relevant I wrote about ePortfolios back in 2006:

    I heard four basic variations on the definitions of ePortfolios at the conference. The first one was the box of papers in the basement. You know, the one with all your notebooks, your tests, your essays…maybe your thesis…? This analogy was introduced by the very first speaker and repeated throughout the day. But the thing is, does anybody ever really think of that box as a portfolio? Personally, I think of it as my “stuff.” If I want to put together a portfolio, I’ll go through my stuff and pull out the best stuff. A portfolio is, roughly, a portable folio. Emphasis on portable. My box of stuff isn’t terribly portable, nor would I have any reason to port it around with me except on those rare and exceptionally distasteful times when I’m moving all of my stuff. I need my box of stuff to put together my portfolio, but the box of stuff is not a portfolio in itself.

    The other three definitions of ePortfolios are closer to the mark:

    1. A periodic browse through the box of stuff: Every once and a while I go down to the basement, pull out my box of stuff, and look through it to remind myself of just how dumb I used to be and how I’ve grown to be slightly less dumb. During those times, I pull out maybe 10% of the stuff in my box. I might pull out slightly different items depending on what I’m thinking about at the time, but it’s always the same process. I pick a few things to read closely and shove the rest back in the box. Reflective ePortfolios should work roughly the same way.
    2. Pulling stuff out to impress somebody: This is the classic portfolio application. When a graphic artist or an architect brings a portfolio to a prospective client or employer, she usually picks a few items from her box of stuff that she thinks will resonate her audience. The collection will be tailored to the particular prospect, just as a cover letter and CV might be customized for each job application. An ePortfolio for potential employers should work the same way.
    3. Pulling stuff out to prove you did the work: Professional eportfolios for certification do this. They collect specific items so that evaluators can easily review the work.

    So to support ePortfolio applications of all types, we need two things: A big box for stuff and some smaller…um…folios that are easy to fill with carefully selected subsets of the stuff. In other words, we need to give students a personal file storage system that’s linked to a personal publishing system. In the former case, the box should automatically store the stuff that students produce or submit online for their coursework. Why let student contributions be “owned” by a course instance which gets archived at the end of the semester, never to be seen again? Why not have it be “owned” by the student and published to the course? Why not have the instructor comments/grades get attached to the document and put in the student’s box, the way comments and grades get attached to physical papers that we return to our students? This isn’t an issue of building an ePortfolio; it’s an issue of correcting a fundamental design flaw in the LMS’s themselves.

    Once every student has a box of stuff, then we can talk about making it easy for them to create portfolios that happen to be “e”. We need a simple publishing system that allows flexible templating and guest access control. Add to the mix a handful of pre-created templates to start the students off, and you’re basically done. You can add bells and whistles–maybe a commenting capability for guests, maybe a simple workflow for reviewers (including the students themselves, in a reflective portfolio application), etc.–but these are all nice-to-have add-ons. They are also, by the way, standard fare for even basic content management systems (like blogs, for example). Let’s keep it simple. An ePortfolio is a lightweight personal publishing system that should sit on top of an LMS’s personal file management system.

    Badges and CLRs should dump into a box of stuff. Learners can add to the box throughout their lives. The technical implementation might be a wallet. However, the user experience must be a box. A wallet isn’t great for organizing lots of disorganized stuff. In any case, this wouldn’t be hard to build. Digital credential wallets exist. In fact, the box-of-credential-stuff product I’m describing probably already exists. I just don’t happen to have seen it yet. I’m not aware of any technical barriers.

    There’s been a lot of talk—and many, many meetings—around the concept of a Learner Employment Record (LER). 1EdTech is involved in some of those conversations, and some of my colleagues are closer to it than I am. I do understand this much: You can’t license or download an LER today. You can’t build one according to a specification. LER is not a thing yet. It’s an idea. I’m less clear on exactly what that idea is. I’ve seen multiple declarations, white papers, and diagrams of LERs from different groups, groups of groups, groups insisting they’re not groups, and groups of groups insisting they’re not groups. Some of my colleagues participate in some of those groups. I sit in when I can. It’s not gelling for me yet; it’s not clear to me that there is a consensus understanding.

    Standards groups, at least in EdTech, are vulnerable to what I call “death by a thousand convenings syndrome”. 1EdTech is far from immune from it, which is one reason I walked away from the meetings during some of the years between when I was contributing to the standards as an Oracle employee and when I accepted my current job under the new leadership of Curtiss Barnes, a person I trust to make things happen.

    I’m a passionate believer in interoperability standards. When done right, they make it economical to deliver real value to users of the tools, make it easier for solving hard and important educational problems in a scalable, financially viable way, and make it harder for companies to profit off of what should be table-stakes functionality (like the ability to ensure student data is handled with appropriate sensitivity or easily add the right educational tools to a particular virtual course environment, for example). But it’s hard to build effective standards coalitions. It’s a Conway’s Law problem. Until you can get a group capable of taking action that’s sufficiently aligned around clearly defined, mutually beneficial standards-making, you’ll see many meetings of disparate stakeholders over multiple years. It’s both a symptom and a cause. This is the litmus test: If you miss six or twelve months’ worth of meetings and you’re not feeling a little lost when you return because of the things that happened while you were away, you probably don’t have the ingredients you need for standards-making in that room. Increasing the ability to recognize and correct that problem is one of the personal contributions I aspire to make at 1EdTech. My sense is that the organization is improving a lot and still has a lot more improvement it can achieve. I apply the same lens to work inside 1EdTech that I apply to work with our coalition partners and across the ecosystem.

    Regarding LER, when I ask folks I respect across the digital credentials world what it is, and I get different answers, that’s a symptom. Maybe an LER is a box of stuff that includes learning-related VCs (e.g., OBs) and employment-related VCs (e.g., a driver’s license). If so, then I’m not sure why it’s complicated. Just create a VC box of stuff and be done with it. I admit I’m neither a standards geek nor a digital credentials geek, so maybe I’m missing some complexity. It’s been known to happen.

    If an LER something different than an expanded box of stuff, then someone needs to explain clearly exactly what it does and how that functionality creates value. Not how it does…whatever the thing is that it does. Unless you’re way down in the technology stack—I’m talking about the level of “make web pages on the internet render properly—nobody is going to rally to the call for an ontology or a transport. They want to know about value. I’m starting to see the coalition-rallying goals crisp up a bit in efforts like AACRAO’s Project Infuse. While I don’t know if Infuse will succeed yet, I do feel like I have a fairly clear idea of what it’s trying to accomplish. And I do feel like I’m in danger of falling behind if I miss a meeting. I’m participating in the governance strand, so I don’t hear the same things that my colleague Rob Coyle hears in the technical strand. But the folks I talk with in the meetings I attend seem to be going somewhere together.

    Likewise, LER-RS, a digital résumé standard being shepherded by HR-Open, makes perfect sense to me. It’s the folio you curate from your box of stuff for a prospective employer. The box of stuff it pulls from aligns well with existing standards. 1EdTech has been supporting HR-Open on this project and is collaborating on a certification suite for it.

    My 1EdTech colleagues who have been working on digital credentials far longer than I have tell me the term LER originally came from a 2020 white paper issued by the US Department of Labor’s American Workforce Policy Advisory Board Digital Infrastructure Working Group. The term was invented to point to a set of functional needs and cited LER technologies that were already in production at the time. I know some of the folks who worked on that paper, and they’re all people I respect. The paper focuses on the “what.” Reading it now, I’m still not seeing any big gaps in the standards needed to make it a reality, at least at my level of understanding. The problem seems to be one of coalition-building. Holding lots of convenings and creating a coalition for action are not the same.

    To my mind, a lot of the LER noise is a side show, not because LER isn’t important as a concept but because many of these conversations do not seem to advance the goal. Meanwhile, digital credentials are advancing.

    AI and the shift

    Regular e-Literate readers know that I try to understand what technologies are good for rather than deciding if they’re “good” or “bad”. AI is a good fit for advancing digital credentials for four reasons. First, it helps on the supply side. A university that doesn’t have defined competencies or the resources to define them can plausibly extract competency descriptions from course catalogs and transcripts. Will the resulting badges and CLRs be great? No. There usually isn’t the right kind of data (like evidence of achievement) in the transcript. Could it be significantly better than nothing? Absolutely. (Again, there is a different potential path, which I’ll unpack in another post.)

    Second, it helps on the demand side. Employers are already having AIs read résumés. Forget about transcripts. A rich, machine-readable, AI-queriable skills record could lower the amount of effort required enough that employers would extract net value from the LER-RS. As a prospective employer, I could ask fairly detailed and sophisticated questions about a candidate pool and have AI surface interesting answers.

    Third, as AI facilitates the evaluation of skill verification assertions, the locus of value in a credential will shift from the issuer to the proof of achievement. A university’s reputation is a proxy for the educational achievement of the student. And it isn’t a great one. While I doubt AIs will be terrific at evaluating a wide range of skill assertions in the near future, they could be good enough to give great students from less prestigious institutions a better chance at getting noticed.

    Finally, AI may help to capture emerging skills that have not yet been codified. For example, recently I’ve been vibe coding as a non-programmer. I’ve figured out how to vibe code Model Context Protocol (MCP) servers in TypeScript and python, use progressive disclosure patterns to reduce AI token usage while increasing accuracy and security, and build a compositor that enables me to orchestrate these workflows using microservices. Some of these skills didn’t exist six months ago. And even if “my” code is good, it wouldn’t tell the story. How did I engineer Claude Code’s context to get it to think like a developer? Did I get it to follow practices that would check its code quality in ways that I can’t, like test-driven development? How did the idea of a “compositor” come about, and how did I make sure it wasn’t over-engineered AI slop? If I did? An AI that understands digital credentials standards could identify, express, and capture evidence for emerging competencies as part of the exhaust stream of my work. And another AI could read that evidence. To be clear, nobody would have any reason to believe that I have any of these skills based on my formal work experience. To bastardize a saying, the proof of the pudding is in the reading.

    When I was hiring Agile Product Owners at Cengage, we used to take the top candidates and run them through simulated product situations to see how they would handle them. We deliberately created complications. Yes, we were looking for craft. But we were also looking for patterns of behavior that are related to how an individual applies a given competency. It’s about how they think as much as what they think (e.g., how they think about the purpose and applications of user stories or retrospectives). What do they bring to the table that’s unusual or unique? It was a time-consuming process involving multiple staff members, but it helped me identify the best performers in a way that no documentation I could have requested at the time would have revealed. If I had access to examples of their real work product, structured in a way that I could interrogate using an AI, I’m not sure if I’d need to run those simulations.

    Still learning

    I don’t pretend to be an expert in digital credentials. Far from it. It’s caught my attention in a way I didn’t expect, though. And I think it’s at just the right level of messiness and foment to be a space where we can make some new and significant progress as a sector.

  • Blursday Socials Are Reborn! First One on Thursday, July 31st

    Blursday Socials Are Reborn! First One on Thursday, July 31st

    CRITICAL UPDATE

    The original post said Tuesday, July 29th. The actual date is Thursday, July 31st at 11:30 AM ET. Sorry for the mistake and the resend.

    For folks who have missed Blursdays, they’re back and better than ever. The first one is coming up fast on Thursday, July 31st at 11:30 AM – 12:30 PM ET.

    Here’s the deal: The Empirical Educator Project (EEP), now under 1EdTech, has been rebranded as 1EdTech Learning Impact Labs. This name change is important because, unlike EEP, 1EdTech can drive change right into the EdTech ecosystem. Accordingly, Blursdays have been renamed 1EdTech Learning Impact Labs Live (or LIL Live). Lots of old friends, some new ones, and a renewed focus on driving change. Our first guest is Unizin CEO Bart Pursel, who has two concrete, actionable proposals for us:

    • We know interleaving works as a teaching practice. The data are clear. Why don’t we create cross-platform interoperability standards that make it easy to implement the same interleaving methods across different LMSs, courseware systems, etc.?
    • We know that course design can significantly impact whether students change their majors. Why don’t we make it easy to analyze the impact of course designs on major changes?

    Bart has already presented these ideas to the 1EdTech community at our Learning Impact conference. Now I’ve asked him back to speak with…well…you. You can weigh in on what you need. You can help shape the 1EdTech community’s perspective on these topics. And you can still enjoy the old Blursday camaraderie.

    How to join

    We’ll be using Engageli. For those who haven’t been to a Blursday and haven’t used the platform before, Engageli is a virtual learning platform designed to foster active learning and engagement in live and asynchronous learning environments. You don’t need to pre-register for Learning Impact Labs Live; just click the link to access the classroom lobby on the scheduled day and time. Sign-up will take a minute or two, so please allow yourself time if you can. Here are the steps:

    • Input your email address and receive a verification code
    • Once you input the verification code, edit your learner name to your first and last name
    • Check your audio and video settings
    • Select Join classroom

    Please feel free to watch this quick video before our session to get a feel for Engageli classroom!

  • EEP at 1EdTech Learning Impact: Solving the Right Problems

    EEP at 1EdTech Learning Impact: Solving the Right Problems

    I write this post to e-Literate readers, Empirical Educator Project (EEP) participants, and 1EdTech members. You should know each other. But you don’t. We should all be working on solving problems together. But we aren’t.

    Not yet, anyway. Now that EEP is part of 1EdTech, I’m writing to ask you to come together at our Learning Impact conference in Indianapolis, the first week in June, to take on this work together.

    1EdTech has the potential to enable a massive learning impact because we have proven that we can change the way the entire EdTech ecosystem works together. (I recently posted a dialogue with Anthropic Claude about this topic.) I highlight the word “potential” because, as a community-driven organization, we only take on the challenges that the community decides to take on together. And the 1EdTech community has not had many e-Literate readers and EEP participants who can help us identify the most impactful challenges we could take on together.

    On the morning of Monday, June 2nd, we’ll have an EEP mini-conference. For those of you who have been to EEP before, the general idea will be familiar but the emphasis will be different. EEP didn’t have a strong engine to drive change. 1EdTech does. So the EEP mini-conference will be a series of talks in which the speakers propose ideas about what the 1EdTech should be working on, based on its learning impact. If you want to come just for the day, you can register for the mini-conference for $350 and participate in the opening events as well. But I invite you to register for the full conference. If you scan the agenda, you’ll see sessions throughout the conference that will interest e-Literate readers and EEP participants.

    EEP will become Learning Impact Labs

    We’re building something bigger. Nesting EEP inside Learning Impact is just a start. Our larger goal is to create an umbrella of educational impact-focused proposals for work that 1EdTech can take on now and a series of exploratory projects for us to understand work that we may want to take on soon. You may recall my AI Learning Design Assistant (ALDA) project, for example. That experiment now lives inside 1EdTech. As a community, we will be working to become more proactive, anticipating needs and opportunities that are directly driven by our collective understanding of what works, what is needed, and what is coming. We will have ideas. But we need yours.

    Come. Join us. If you’ve been a fellow traveler with me but haven’t seen a place for you at 1EdTech, I want you to know we have a seat with your name on it. If you’re a 1EdTech member who has colleagues more focused on the education (or the EdTech product design) side, let them know they can have a voice in 1EdTech.

    Let us, finally, raise the barn together.

    Come.