Jimmy Wales at SIMS — transcript
Machine-generated transcript — not human-reviewed
About this transcript
The original audio is authoritative. This draft may mishear words or names, omit speech, and merge questions with answers. Changing voices have not been identified, and meaningful non-speech sounds have not been described. Unflagged text may also be inaccurate.
The full recording was submitted to local speech recognition. Generated text is preserved in its original order, including suspected repetitions. Machine-flagged uncertainty identifies possible errors, not verified corrections. Gap notes may indicate silence or omitted speech; they do not establish which.
Download the original recording (MP3 audio, 1:17:33). Timestamp links open that recording at the indicated time where the browser supports audio fragments; the displayed times can also be used to seek manually.
Send a transcript correction to Joseph, including the timestamp and the words you heard.
Generated transcript
0:00–5:00
So thank you all for coming. For those who aren't from here, this is the School of Information. I brought Jimmy Wales, who is the founder of Wikipedia, here today because I think it's very relevant to all of our issues around information. I'm going to let him do whatever other directions he thinks is appropriate. Thank you. She said the reason she brought me is because it's relevant, but I think she really just wanted to give us some exercise. Come on up that hill. So, yes, I'm Jimmy Wales. I'm the president of the Wikimedia Foundation and the founder of Wikipedia.
And today the main thing I'm going to talk about is a basic overview of Wikipedia, how the community works, the core principles of the Wikimedia Foundation, and what will be free. I have a list of 10 things that I think will be free in the future, and it's basically my manifesto I've been working on, so you'll get to hear all that. So in 1962, Charles Van Doren, who was later a senior editor at Britannica, said,
the ideal encyclopedia should be radical. It should stop being safe. But if you know the history of Britannica since 1962, it's really been anything but radical. It's still a very safe, boring, old-fashioned encyclopedia. Wikipedia, on the other hand, begins with a very radical idea. And that's for all of us to imagine a world in which every single person is given free access to the sum of all human knowledge. And that's what we're doing. So the Wikimedia Foundation is our non-profit organization that I founded. The aim of the foundation, of course, is to distribute that free encyclopedia to every single person on the planet in their own language.
The foundation is responsible for Wikipedia and all of our sister projects, which I'll tell you about a little bit. We're funded primarily by donations from the public. We have fund drives on the site from time to time and people donate money. We're also partnering with select institutions. So for example Yahoo donated some servers for use in our South Korea facility and Kinesnet, a Dutch educational consortium, gave us some servers for our facility in Amsterdam. So what is Wikipedia? First of all, how many people here have used Wikipedia?
I expect. Is there anyone here who hasn't? Okay, leave now. How many people have actually edited Wikipedia? Okay, so also a very good number. How many people here would call themselves Wikipedians? How many people have donated? Not yet. Not yet. Okay, so we've got one or two Wikipedians here. So Wikipedia is a freely licensed encyclopedia. It's written by thousands of volunteers in many languages. The really, every part of that statement carries some weight.
So freely licensed is very, very important. What do I mean by free? I mean, as the free software people, I mean free as in speech, not free as in beer. So many years ago Richard Stallman formulated what he calls the four fundamental freedoms of free software. Those freedoms are the freedom to copy, the freedom to modify, the freedom to redistribute, and the freedom to redistribute modified versions. And to be able to do this commercially or non-commercially. So those are the four freedoms that make up the cornerstone of what's called free software.
So GNU, Linux, Apache, all that stuff that really runs the internet is all free software. and every part of Wikipedia is free as well. It's an encyclopedia, so we make a distinction. Wikipedia is not a joke book. Wikipedia is not basically a place for dumping all kinds of information. It's a fairly standard conception of what an encyclopedia should be, what the entries are, encyclopedian entries. It's written by thousands of volunteers and then in many languages,
and I'll tell you more about the languages in a minute. So how big is Wikipedia? The English Wikipedia is the largest and has well over 500 million words. Since our developers are all volunteers, I have a hard time getting them to bother with calculating such statistics for me because the only real reason we need to know is propaganda purposes. So I believe that based on the number of articles we have when they calculate the 500 million, that we're getting very close to the 1 billion word mark in the English Wikipedia. English Wikipedia is larger than Britannica and Encarta combined.
The German version of Wikipedia is equal in size now to Brockhaus. Brockhaus is the German equivalent of Britannica, so it's the standard high quality old encyclopedia. So how big is Wikipedia globally? Our project is fully global. It's important to understand that only about one third of our total work is in English, and only about one third of our total traffic is to the English language Wikipedia. So we're talking about a project that has a very diverse group of people from all around the world. We have over 750,000 articles in English.
5:00–10:00
That should be 800,000 soon. It's not today. I'm not sure exactly. I forgot to look this morning. We have now over 300,000 in German. We have 100,000 each in French, Japanese, Polish, Italian, and Swedish. So those are the ones who've cracked the 100,000 number. And over 50,000 in Dutch, Spanish, and Portuguese. All total, we've got more than 2.2 million across 200 languages. But I feel that the 200 language number is not the right number to use. 200 languages, that's the number of websites that we have set up.
But a lot of those websites are actually just, they're sitting there waiting for someone to even come and translate the interface. It just says, hi, welcome to the whatever Wikipedia. If you know this language, please help us translate the interface. So that doesn't really count. The real count would be that we have 30 language editions that have at least 10,000 articles, and we have 75 language editions that have at least 1,000 articles. So 1,000 articles isn't really much of an encyclopedia, but that's the point at which I would say there is a small community there, there's usually four, five, six people who are regulars who are working on it,
and they're starting to get enough content that it begins to attract people. So that's what I would consider 75 active projects. So we have several other projects under the umbrella of the Wikimedia Foundation. In addition to Wikipedia, we have projects which most of these projects came about because of some kind of social pressure within the community. The first example would be Wiktionary. Wiktionary was created because people started putting dictionary type information into the encyclopedia.
So, for example, etymologies, antonyms, synonyms, all that kind of thing. And we realized, well, really, a dictionary is a fundamentally different type of reference work from an encyclopedia. And so it made sense to have a separate project for all the people who wanted to work on the dictionary stuff. Wikibooks, this is the project that I'm personally always most excited about for the long run, because I think it really will be our most important legacy. WikiBooks is an effort to create freely licensed textbooks for kindergarten all the way through
the university level. This is important for our big picture mission. When we say we want to give a free encyclopedia to everyone on the planet, we don't mean we're planning to spam them with AOL-type CDs that they can't even use. We want to give them an encyclopedia that they can actually use. So implicit in that is they need to have all of the literacy materials that they would need to be able to come up to the level to do that. So right now in wiki books there are already some 10,000 modules as they call them underway
and more and more people all the time working on it. So I think that's going to be very exciting in the future. WikiQuote is like Bartlett's familiar quotations. Again this grew from some social pressures within the community. People were putting way too many quotes from famous people into the encyclopedia articles and then people would get into editorial fights about it. So in order to relieve that social pressure, we said, well, all of you who really enjoy just going out and finding quotes and validating them and getting the site, please put those over in WikiQuote. And then an encyclopedia article can, of course, legitimately have a few famous quotes from someone.
So that has worked out very well. Wikimedia Commons. This is where we put all of our media files. So this would be images, sounds, and video, although we don't have a lot of video and we don't have a lot of sounds. We have in Wikimedia Commons now over 200,000 different media objects, mostly pictures. Wikimedia Commons came about because before we had the Commons area, people would just simply upload photos to their own local language Wikipedia.
So if you were working in the English Wikipedia and you wanted a picture of the Eiffel Tower, you could probably guess that you could probably go and look in the French Wikipedia and find it and you'd be right. On the other hand, if you wanted a picture of something in Thailand, it probably wouldn't occur to you to go and look in the Dutch Wikipedia. But in fact, that's where there were a lot of pictures of things in Thailand because we have a very prominent Dutch contributor who happens to live in Thailand. So we decided that, well, really, these media files are generally language neutral. They need to be gathered in one place so that they can be more easily found by everyone.
The other thing that we have to deal with, it's a significant issue for us that we're always discussing and thinking about is how to comply with the law in different jurisdictions. So since the servers are in the United States, or the bulk of the servers are in the United States, and since the foundation is in the United States, we have to comply with U.S. law. That's actually one of the easiest jurisdictions for us to comply with because the U.S. has very generous fair use provisions. Under the Digital Millennium Copyright Act, we have protections.
If a user uploads some copyrighted thing and we get a complaint, We just have to take it down and we can't be sued for that. On the other hand, we also have an interest in, as far as we can, also obeying the laws of other countries. So for example, the German language Wikipedia doesn't allow fair use images because under German law, the concept of fair use is essentially non-existent. So therefore, in the German Wikipedia, they follow both US law, but they also follow some additional restrictions that would be German law.
10:00–15:00
This applies really only to copyright considerations. We don't follow local law if it comes down to a really serious censorship issue. So that really doesn't happen all that much, but there are potentially some issues there as well. WikiNews is our newest really big project that went through a full process of community approval. WikiNews is in five or six languages now. but really the only really active Wiki news projects are in English and German.
This effort grew out of, we noticed that anytime there's any sort of a major breaking news event anywhere in the world, the Wikipedia article tends to be very, very good. It's a nice summary, synthesis of all different kinds of news sources and incredibly detailed. And then also people come in and they fill out a lot of background information. So when the tsunami happened, you could turn on CNN 24 hours a day and see people's vacation movies and, you know, ah, water, you know.
But if you really wanted to learn who are the people who are affected, what are their lives like, what is their government like, what language do they speak, all that kind of basic background information, Wikipedia was a great place to turn because all that information got very filled in very, very quickly. So having seen that kind of thing, we thought, well, you know, this wiki thing could be useful for news in general and again we had a group of contributors who were always working on in the news and current event topics who wanted their own space to work so we moved that into WikiNews and that's doing reasonably well.
So how popular is Wikipedia? I guess it's very popular here at Berkeley, it seems like everyone uses it, but we're now a top 40 website and it's actually top 30 depending on the day of the week according to Alexa.com. And we have a broader reach, by reach I mean the number of unique visitors to the site, the number of people who are seeing our site in any given month. We have a broader reach than the New York Times, broader than the LA Times, Wall Street Journal, MSNBC.com,
Chicago Tribune. But the interesting thing about this is that we have a broader reach than all of those outlets combined. So we're seeing more people see Wikipedia in a day than seeing all of those things combined. So I really like this when journalists ask me, they call me up and ask me some question about the mainstream media, and I say, you mean Wikipedia? Because it's become very popular. So we're doing around 2.4 billion page views a month. This number is actually a little bit out of date.
And this graph shows the history of our growth. and here when we went, just before we passed the red line, the red line is about.com which was sold to the New York Times for 410 million dollars. So we really got a big kick out of this because we're like this crazy bunch of volunteers on the internet and we've managed to build something that's of huge economic value. That's one of the most important things to know about the Wikimedia Foundation is that we are a very, very tiny nonprofit organization.
So we have exactly two employees. We have our lead software developer and we actually, when we hired Brian, it wasn't because we needed him to work more for Wikipedia. It was that he had, for two years, he had worked a part-time job 20 hours a week and worked way more than full-time working on the software and everybody thought it would be good if Brian got a life. So we hired him so that he could go to the movies sometimes instead instead of just working all the time. I don't think it's really worked, but we'll see. So, our hardware, now we have over 120 servers
in multiple data centers, so we've got servers, the bulk of the servers are in the United States, in Florida, but we also have servers in Paris, and in Amsterdam, and in South Korea. So we've got all these data centers, and they're all managed by volunteers. So I don't think any company could really replicate the kind of infrastructure that we've got because right now I could go online, go into an IRC channel and there would be 8 to 12 very qualified website engineers who know the site inside and out
and they're monitoring everything all the time. And it's completely organized in the wiki way. We don't have schedules, we don't have hierarchy. It's just whoever shows up, whoever's online, whenever they're online, keeps an eye on the servers. And we have enough people that that works out reasonably well. Of course, they have each other's phone numbers, so in case of an emergency, they can call each other and wake them up. They no longer call and wake me up because I'm useless. So they learn that pretty quickly. So this is the layout of our network.
15:00–20:00
I don't intend for this to be a technical talk, but I do think it's kind of worthwhile for people to get an idea of what is going on here. The Internet's out there, so you're out there on the Internet, and you're going to request a page. and you request it from this group of servers here, which are called squids. Squids are servers which cache webpages, so if they've already seen the page, they just hold it in memory in the hopes that somebody else is going to ask for the same page. And in this picture, this is an old picture, you may get the page from the Florida squids or from the French squids. This is when we had only two data centers.
Now you'll get them from whatever squid is geographically closest to you. So from there, if the squids don't know the page, then they have to compute the page. And so then they go to the Apaches. And so this is a huge farm of Apache rendering machines. They basically get information from the database cluster over there, and they build a web page out of it, and then they send you the web page. There's a couple of other servers for image storage, mail servers, DNS, all those kinds of things. This is a very standard type of web architecture.
And it's very scalable. In order to handle more traffic, all we need to do is just add more of everything here. Because it's just more and more of everything. So in terms of how Wikipedia works, there are basically two views of Wikipedia. The first view is that Wikipedia is an emergent phenomenon. Something like pseudo-Durwinian. This is where you hear phrases like collective intelligence, hive mind, all those kinds of things.
The other view would basically focus on the community of thoughtful users. So a good introduction to the quasi-Darwinian model comes from a former Britannica editor who's most famous for comparing Wikipedia to a public toilet. But he also wrote something a little more intelligent, and he said, some unspecified quasi-Darwinian process will assure that those writings and editings by contributors of greatest expertise will survive. Articles will eventually reach a steady state that corresponds to the highest degree of accuracy.
Does someone actually believe this? Evidently so. So when I first read this, I thought, you know, that's really interesting, because very rarely I do hear people kind of talk about Wikipedia in this way. But within the community, within this core group of 600 to 1,000 users that I talk to all the time, we don't think about Wikipedia in that way. This view, the emergent phenomenon view, would suggest that we're like ants. That there's thousands and thousands of individual users, they don't know each other, they each contribute a little bit, and somehow out of this emerges a coherent body of work.
The other view, and you can tell from my slide here, that I have a bias here, because I put pictures of a bunch of my friends up there, that we're a community, a dedicated group of a few hundred volunteers who know each other and work to guarantee the quality and integrity of the content. So those are the two views, and when I was thinking about this, I thought there really are some implications for these two views in terms of how I manage the project. The goal that I've set for myself in life is to distribute a free encyclopedia to every single person on the planet. I really could care less about wikis.
That's our method of getting our work done. And so there are things in my management of the project that would flow from either of these views being true. In the emergent model, you need reputation mechanisms, things like what eBay has and Slash. has. The reason eBay has their reputation metric is precisely that eBay is not itself a community. Now there are some communities within eBay, but eBay taken as a whole is not a community. The reason when I go to eBay I care about somebody's reputation metric is I don't know the person,
and I don't know anybody who knows the person, I really can't ask around. Instead, all I can do is rely on this number. And so if in the emergent model where there's thousands of people who don't know each other, you need some kind of numerical mechanism to rate people or contributions. It also implies that individual users are tiny and have no power. We do a lot of struggling within the community about due process in terms of blocking people who are being a pain in the neck. If it turns out they're all like ants, well, you know, you can step on a few ants. it doesn't really hurt the colony, so maybe we could just do that and not worry about it.
The community model, on the other hand, suggests that reputation is a natural outgrowth of human interactions, that instead of having a single number that you rate people by, instead what you have is human judgments. Human judgments would say, this person does really great work on biology articles, but don't let him go anywhere near Israel-Palestine because he goes berserk. You know, I have things like, this person does really good work, but they're very rude to other people. So how do we deal with that? Those kinds of human judgments are exactly the kind of judgments that you make within any type of organization.
20:00–25:00
So say within a university faculty, within a church group, within a company, that you don't go around having reputation metrics inside those kinds of organizations, because you know people and you don't really need that. And humans are quite good at holding in our minds this very complex, layered notion of relationships with others and their reputations. And then the other implication of this is that all these users, the particular users, these power users who are around me all the time, and they're on the mailing list, and they're in IRC, and they're on the website,
if they're the ones actually doing the work, then they actually are powerful. They have to be respected, because if we don't respect their wishes, they're liable to stop doing all the work. So those are the things that I looked at as being the implications. And so I did a study looking at the edit history of Wikipedia. And since I did that study, the developers changed our statistics routines so that these kind of numbers are gathered all the time automatically. So that's very convenient for me. What I expected to see is something like an 80-20 rule.
80% of the work being done by 20% of the users. It turns out the distribution is actually much, much tighter than that. About half of all the edits are done by just 0.7% of all users. That's 615 people. That's 615 people that have been responsible for half of all the edits to English Wikipedia. This number is a couple months old, so it's probably a few more people than that now. The most active, let's call it 2%, have done almost three-fourths of the work.
So this is 1,700, 1,800 people are responsible for three-fourths of the work. So what this says is that this core community, they're the ones who are actually writing the site. And we do have edits by anonymous users. This is controversial and intriguing. Yes, you can edit this page. You can go to almost any page of Wikipedia except for the front page and a handful of other pages, very high-profile pages. and you can click edit, make a change and save it.
You don't even have to log in. This is one of the first things most people learn about Wikipedia and it's very mind-boggling. But it turns out that although anonymous IP numbers can edit Wikipedia and they do, those edits make up only about 18% of all the edits and there's some evidence of a downward trend over time. It used to be 22%, then 20% and then the last time I checked it was around 18%. And then anecdotally, many regular users report sometimes editing anonymously by accident. That's me because I always forget to log in. Or as a quiet form of sock puppeting.
So sock puppeting is what we mean by you've got one account and then you've got someone over here. That's your sock puppet. Sock puppeting, by the way, is mostly used as a pejorative term, but not always. It's possible there are many legitimate reasons a person might have two accounts in Wikipedia. One example, someone wrote to me and said, I'm doing a lot of work in Wikipedia in my area of professional expertise and a lot of my colleagues at the university know I do this. And I assume they check my contributions from time to time or poke around in what I'm doing in Wikipedia because I'm always talking about it.
But I went to the article on pedophilia and it's really, really bad. It's just wrong. And I want to edit it, but I don't want my friends to come and see that I spend my days working in the pedophilia area. They might get the wrong idea. So I want to know if it's okay if I have a second account. And I said, well, of course. You don't even have to ask me. Just do that. The things that people use a second account for that is socially completely unacceptable would be things like voting twice in any kind of thing, or getting into an argument with yourself, or chiming in to support yourself.
Oh, believe me. We see it all. And it's actually a great way to win an argument. You know, you start arguing with people, and then you get some detractor who's just like a moron, right? Everybody's like, wow, okay, I see you're right after all. We have ways of checking for sock puppets. We have a check user tool. And it's actually one of the interesting things about the way our community culture works is that we're very, very cautious about who has access to the check user tool
because we do value our users' privacy. And if people have two accounts and they've got their own personal reasons for that, we really don't want to know it. Yet we can dig into the database, we can dig into the archives of the access logs and figure it out, but mostly we don't. First of all, it's kind of a lot of work, so it's just there's so much data flowing in. The other thing is we don't really want to know. The only time that we want to do checks for sock puppeting is if there's an actual problem on the site that we need to check. So we control access to that pretty carefully. So how does this community ensure quality?
How does the software empower good work? I mean, the main thing that most people find when they learn how Wikipedia works is the first thing you think when you hear there's an encyclopedia and anyone can edit it, you just think it must be complete and total crap. That's completely impossible and insane way to run a website. But then you start looking at Wikipedia and you find out, you know, it's actually not so bad. In parts it's actually very good and in parts it's, well, it could be better, but it's actually surprising that it reads reasonably well in the most part.
25:00–30:00
The facts are basically okay. It isn't like a lot of things you find on the internet where you get really one-sided rants. It's mostly neutral. And so the question is, how do we do this? How does this wide open editing model work to actually ensure quality? So the most important thing to know is the idea of real-time peer review. The community controls the quality of work by monitoring everything that goes on on the site. So every edit goes onto the recent changes page, and the recent changes page is watched by hundreds of people daily.
It's actually interesting to know that in English and German and French and Japanese, really the largest Wikipedias, the recent changes goes by so fast it's basically not directly usable anymore. And so what has happened within the community, all of our recent changes are pumped into an IRC channel. so that happens automatically and then people have built tools which listen in on the IRC channel and actually categorize the edits in various ways so you can whitelist users you can blacklist users
you can look for radical changes of size you know if something went from 20k to just 8 characters it's probably F-U-C-K space Y-O-U that's one of the more popular ones so people are monitoring that sort of thing and they're building tools as the site grows to make it easier to monitor those things. Additionally, every user can set up their own personal watch list. Most active users have this set up as their default that anything they edit automatically
is added to their watch list. So then every day when you go to the site you can just click on my watch list and instead of seeing all of recent changes, you just see the recent changes in the area that you're interested in. So, I actually met a Cornell University ornithology professor who told me he keeps a bunch of bird articles on his watch list and he comes in maybe only once a week but he checks all the edits and that's just his little bit of community service to see how it's going in the area of his expertise. There are many other things in the software like the new pages tool so we can immediately
see when new pages are created in order to determine if they need to be deleted or something like that. So the page history, there are a lot of different wiki software engines out there. I used to be able to confidently say we had the best diff page showing the difference between the old version and the new version. I'm not really sure that's true anymore because I assume some of the others have evolved. There used to be a couple that actually just gave you the output of the Unix diff command, which is really only readable by Unix geeks and things like that.
So you can see here that the words that are changed are highlighted in red. The paragraphs that have changed have different colors. This means that when you think about what it would take to monitor an article over time, it doesn't mean you have to reread the article from scratch every time, in which case you might miss some little word or phrase. You can actually just do a diff between any two versions that are in the version history. So we keep every single version of every single article. The only exceptions to that is that occasionally we have to delete particular revisions for legal reasons.
So for example, if it's a copyright violation and someone's complained about it, or if it's libel that someone's complained about, then we go in and we have to delete that revision. But by and large, every single version of every single article is there. And so if an article has started to go downhill, anyone can come in and simply go back to an old version that was better and restore it to be the front. So organization by the community is also very important.
The free-form nature of the wiki software lets the community determine for itself how it wants to interact. So the example I'm going to talk about, and there are many, many examples within Wikipedia, is the Votes for Deletion page, which was just recently renamed to Articles for Deletion for very complicated reasons, which you'll understand in just a minute. This particular example here I have is a twisted issue. So this was supposedly a film from 1988. Somebody says supposedly an underground punk film but miserably fails the Google test.
The Google test is you search in Google and if it's not in Google it probably doesn't exist. So the Google test is very widely used in Wikipedia but it's not definitive or it doesn't necessarily prove it but it's a pretty good indicator. If you can't find it in Google you pretty much need to justify where it comes from. So suggesting the famous alpha. So delete, delete. And so the delete, delete here, this is just a wiki page. People are just editing and adding their comments and they're voting either delete or keep.
30:00–35:00
So then the next user says, tentative week, keep. But then later he crossed out tentative week and he says keep. And he found it. He says, wait, I found it in the Film Threat Video Guide to 20 Underground Films You Must See. We should clean it up. And somebody else says clean it up. And then finally Rick K says, keep it, it's a real movie. I found it in IMDB. So then keep, keep. So the result of this was a decision to keep this article. Now the interesting thing about this is that although we call it votes for deletion and people are voting, and although the administrators who go through have a rule of thumb they use, 75 to 80% support for deletion, get something deleted,
it is not in the software. So we very often have programmers who come and propose things like, well, the votes for deletion page is really long and complicated and it's a lot of work to maintain. What we really should do is automate the process. So then they usually have like an eight page summary of exactly how they think it should work. You know, we should have a two week period and at the end of the two week period the vote, you know, if it's more than this or less than that then it's automatically deleted otherwise it could da da da da da da da. And we just say no, no, no. We're going to do it this way because this is the way that preserves the human dialogue
in the community. This way avoids a lot of potential gaming of the system. If you have software that imposes the rules on the community then people want to do something else so they'll start to game the system by sock buffeting, voting multiple times, all those kinds of things you have to deal with. And then additionally, we see here that this is a dialogue, it's a discourse on the merits of the article and an administrator is perfectly free to go through and see the last three keeps and realize that well all the people who voted above didn't have access to the
same information. And I know who they are and I know they're reasonable so I know if they had seen that it was in IMDB, they would have voted otherwise. And so this vote could have easily been 20 to 3 and we would have still probably kept the article. So Wikipedia governance. Wikipedia governance, so how this community is actually governed, is a very confusing but workable mix of consensus. And by consensus what I mean is that we strongly discourage people from voting on the content
of the articles. The reason for that is if you look around the room and you see that you've got 70% support for a particular view It can be very tempting to simply dig in your heels and marginalize the minority view and say well the 30% doesn't matter They're never going to win the vote And that's really not good for neutrality. It's not good for completeness Instead what we do is we encourage people in a situation like that to keep rewriting keep working try to find a way to bring in Because if 30% of the community is dissenting and they're reasonable people that's a pretty major dissent
Now, we don't mean unanimity. You can't always please everyone, because some people are just crazy. And so there's nothing to be done about that, so you can't really have that strong of a standard. But consensus is basically this idea, we need to keep working on it until pretty much everybody agrees. We do have some democracy, by that I mean voting. An example of voting you just saw, votes for deletion, but as you saw there, it isn't rigidly enforced. It isn't built into the software.
It's an open-ended kind of voting that just depends on human judgment. Other cases of voting that we might have, suppose you've got the article on Paris, and you want to have a picture there, and you want to have a picture of the Eiffel Tower, but there's two different competing pictures of the Eiffel Tower that could go at the top of the article. and aesthetically some people prefer one to the other. It's an either or decision. You don't have room for both.
There's no way to edit the two together. That doesn't really work. So you have to choose one over the other. In a case like that, people do hold non-binding polls, but the purpose of those polls is actually to build consensus because most people would say, well, I prefer A to B or I prefer B to A. However, if there's a vote and 65% go one way or the other, then I'll defer to the judgment of the majority and that's fine. So when people are reasonable, voting is actually a good way to gain consensus. And then in such a case, even the people who preferred B would probably revert the change if some very persistent person kept switching it back.
So that's a way of getting some kind of peace and resolution to an issue in the community. There's a certain amount of aristocracy in the community. So an example of this would be, you know, I mentioned Rick Kay on the previous, on the Votes for Deletion example. Rick Kay is a well-known user. He's been around for a long time. He's very experienced on votes for deletion. So we know he has, he knows all the precedents, he knows sort of what's going on. Completely trustworthy, so if he says he found an IMDB, you don't even need to click on the link.
Of course he found it in IMDB if he says so. And so, in a lot of cases, a user like that will carry significantly more weight than someone who we've never heard of. So you don't know someone, they come in and they make a comment. that comment has a stand or fall on its own merits without really a personal reputation behind it. So the aristocracy component is actually quite important. Another example of this would be Angela. Angela was voted from the community to be on the board of the Wikimedia Foundation
35:00–40:00
with more than twice the number of votes of the person who wasn't elected. And she's by far the most popular, most powerful user for a long time within the English Wikipedia. And I always say that Angela could break any rule of English Wikipedia and get away with it because she's that powerful. But on the other hand, it's because Angela's the one person who you know would never, ever, ever, ever break any rule of Wikipedia. And I like to tease her that she's the only person who actually knows all the rules of Wikipedia.
So this aristocracy is an important part of how the community works, that there is an elite group of users, very informally identified, but who carry a lot of weight and actually have a lot of control. And then finally there's monarchy, and that's my role in the community. I've given this talk in Germany, and the next day in the paper, the headline said, I am the queen of England. which is not exactly what I said, but the point of the analogy is that if you're familiar with the workings of the free software movement
in those communities, there's a very long tradition of the benevolent dictator model. So you have somebody like Linus Torvalds, somebody like Larry Wall and Pearl. These are people who are considered to be basically the benevolent dictator of their little area of the free software world. And the reason for that is not that programmers are prone to tyranny, it's that if you're trying to organize a group of people, a small group of people, along the lines of rough
consensus and running code, that's an old saying, and you don't want to get bogged down in a lot of formal decision making mechanisms, which can of course yield the wrong answer. So if you're voting on everything, you can have contradictory votes and things like this can happen. Instead, it's very useful to just say, this is the person who's in charge and will agree to defer to their judgment. That person has to be the right type of person. The successful leaders of free software projects tend to be a certain type of person who listens well to others, tries to work for compromise, tries to get everybody on board.
Those are the important qualities that allow somebody to do that. Well, I don't think it's appropriate. I don't think it's ethically appropriate for anyone, especially not me, to be the benevolent dictator of all human knowledge. That's a bit much. So at the same time, in the very early days of the project, of course, we needed to move forward with things. And so we had this sort of a benevolent dictator model. We're moving away from that. And this is where the analogy of the British monarchy comes in. that as we grow institutions in the community to handle the kinds of functions that a benevolent dictator would handle,
then we're able to take those functions off of me and trust in the stability of the community to take care of things. The reason we don't immediately jump to a democratic model has to do with exactly that we're a fast-growing community. It's very experimental. One of the things, I can give an example of why we still preserve a certain amount of power for me, We had a case where a neo-Nazi website discovered Wikipedia. So they had on their message boards, they found this horrible site, Wikipedia, which is clearly a Jewish conspiracy, blah, blah, blah.
And they said, well, we're going to go in. Somebody said, hey, look, I found they vote to delete things. So we're going to go in. There's 40,000 of us, and we're going to storm in, and we're going to take over Wikipedia. So they did. They stormed in to take over Wikipedia, all 18 of them. That's Nazi math. So 18 people came and voted to delete an article that clearly any rational person would say, these are Nazis deleting an article they don't like. So the vote in that particular case was something like 85 to 18. They really had no chance of actually succeeding.
But certain people in the community were very concerned. They said, well, obviously these are crazies and so on. What are we going to do if they get more organized? What are we going to do if a group actually does come in and try to take over? And I said, this is easy. We have our rules. We have our procedures, but we made them up. We can change them. And so that's my role in the community is I'll just ban them. We'll just take care of that. That's fine. It's no problem. So that I stand in the role of defending the community to make sure that our democratic processes don't kind of go off track and send us down the wrong path,
that I'm there as kind of a check on things so that the core community can say, hey, wait a minute, wait a minute, wait a minute. Something's gone wrong here. We have to figure this out. We can actually change the rules. We're actually going to have that very soon because last year we had elections to the English Arbitration Committee, which ended up being very, very bitter and very divisive. And after a lot of thinking and discussing and so on and so forth, we're going to still have some form of election, but we're going to actually change the structure of it. And that's just a decision that I've made that everybody's going to be very relieved about because it was such a fiasco last year,
40:00–45:00
that we want to have a situation where, yes, the elections are a means for the community to give input and feedback on who should be on the arbitration committee, but not be a magnet for the trolls to come in and just raise hell with everybody. So we're subtly changing the structure of that to make it a little bit more healthy. So the final point here is that Wikipedians are very flexible about our social methodology, that we value the results over the process. And I think that's really, really important to our success. if we had very early on adopted some sort of pure two-thirds majority voting rule for voting on everything,
I think we wouldn't have gotten as far as we have because we're now able to use different kinds of decision-making procedures for different kinds of things as human judgment would tell us that we need to. So that's essentially all about the community. And then I wanted to talk about the core principles of the Wikimedia Foundation. One of our core principles is free knowledge. that everything we do, all of the content, has to be under a free license. So the bulk of our work is under the GNU free documentation license.
A lot of the images are under the FDL, plus they're also under some Creative Commons licenses or their public domain or whatever. This is really important for several reasons. One reason is that it decreases the individual sense of ownership of the content and it increases the shared sense of ownership, a sense of shared ownership. A lot of collaborative writing projects have fallen down on the problems of ego. Somebody writes something and they really don't want other people messing with it. This is my essay and you really can't get beyond that.
Well, if everything is freely licensed and you understand when you're submitting it that people can modify it, redistribute it, copy it, they can do all kinds of things with it, it allows you to kind of let go and you say, well, it's not mine, it's just some stuff that I did on Wikipedia. But it also increases our sense of shared ownership. The people within the community really take a love for the project and a care for the quality and really making sure that the vandals are blocked and that the site is maintained and things like that. Another perhaps surprising part of what the license does is that it enhances the popularity of Wikipedia.
And the attribution requirement in the license extends our brand name recognition. So there are over 200 websites out there which are simply straight clones of Wikipedia. They just copy the database, they dump it, they throw up some ads. and that's it. That's what they did. But it turns out that none of those has even a tiny fraction of the traffic that Wikipedia does and every single one of those links back to us. So we get traffic from that in basically three different ways. We get traffic, direct traffic from those links. We get traffic from the search engines.
So for example, Google, to the extent those other sites are popular, it gives us a boost in Google because Google looks at who else is linking to you. And then finally, it just puts our brand name out there. So people may go to this very annoying site with lots of flashing banner ads and things like that. They read the article and the bottom says this comes from Wikipedia. And they then go, jeez, I think I'll go to Wikipedia next time because it's much nicer. So I think that a huge part of our growth has been because of the free license. And that's very counterintuitive to a lot of people who think that in order to have a popular website,
you need to have some unique proprietary content that no one else has. If you have this great content, then people will come. The internet just simply does not work that way. That is a completely non-functional model of how to have a popular website on the internet. What really does work is openness. Put the information out there, let people copy it, let people do whatever they want with it, and that brings in directly and indirectly tons of traffic. Another core principle that we have is our neutral point of view policy. We call this NPOV in the community, and we specifically use it as a term of art.
It's jargon within the community because we want to distance ourselves from such words as non-bias, objectivity, truth. And the reason for that is that NPOV is our social concept of cooperation. Wikipedians come from very diverse political, religious, cultural backgrounds, increasing diversity all the time. And we need to find a way to be able to work together in relative harmony. Simply saying that everything should be true doesn't really get you very far.
Two people may not agree on what is true. but it turns out they actually can agree on a neutral presentation of the dispute itself. So the simplest example I like to give is some topic like abortion where people have very divergent opinions. However, even a priest and someone from Planned Parenthood, if they're reasonable, thoughtful people, they can work together on that article because they're able to say the position of the Catholic Church is this and the Pope has said that and Planned Parenthood has said this and that And they're able to actually present that in a way that they both would say,
45:00–50:00
yeah, well, if somebody new came up to us and wanted to know what it is we're arguing about, this article would give them a good, reasonable introduction to the topic. So the neutrality policy basically says Wikipedia shouldn't take a stand on any controversial issue. Instead, we should find a way to present it that's acceptable to a consensus. And then finally, a core principle is free software. Every single piece of software on the website is free software. So MediaWiki is our Wiki engine and it's under the GPL.
And everything, it's GNU, Linux, Apache, MySQL, PHP. It's all your standard free software that runs everything. The reason for that is it's basically ideological. It's the idea that if we were to supply people with a free encyclopedia, not only does the content need to be free, but all the tools that you would need to be able to use the content also need to be free. And so we stick to that rule very, very firmly. So finally, I'm going to go through my 10 things.
How much time do we have left? You have including questions. Okay. Until 5.30. Okay, we've got time. I'll just do this very quickly. So in 1900, the mathematician David Hilbert posed 23 problems, which subsequently became very influential in mathematics. The first ten were announced at the International Conference of Mathematics. Then over time he added some to the list and then he published a paper. And then these problems were very influential. Of the 23 problems, a lot of these, they drove a lot of the agenda of mathematics over the coming century.
and most of them have either been solved by now or they've been declared insolvable or too ambiguous to really figure out exactly what it's supposed to be or what the answer would be. And I thought, well, I'm completely arrogant, so I'll do the same thing for the free culture movement. And so I came up with my problems. These are ten challenges for the free culture movement. And these are designed to spark and inspire innovative work. I originally wanted to call the list 10 things that must be free, but I felt that must
implied either some sort of sense of inevitability That's a little too strong or it might sound like kind of an impotent political demand You know we demand these things to be free. I don't really mean it that way. I mean these are things that I think we can do If people will just recognize that we can't do them and let's just go do them now None of these ideas are unique or original to me. I only list them to focus attention So a lot of these are actually already underway in different places. So the first one, free the encyclopedia. This is Wikipedia.
In English and German, just to pick an arbitrary cutoff point, if you have a broadband internet connection, our mission is done. You've got your free encyclopedia. It's there. You can download it. You can read it. We've distributed it to you. Wikipedia needs a lot of work, but it's reasonably high quality. And so fine, it's done. You've got your free encyclopedia. French, Japanese, etc. are not far behind. The number I'm using for a cutoff here to say we're done is 250,000 articles for no particular reason. But there's a ton of work left to do globally.
Although we're very strong in the languages that I mentioned, we're very weak in a lot of the languages that you might expect we'd be weak in. So like Swahili, for example, is a very, very small project. So there's a ton of work left to do to get things to people in their own language. Free the dictionary, this is another project of ours, Wiktionary. Not as far along as Wikipedia, it's picking up steam. Here we need software development for support. So if you think about the nature of dictionary data,
it's very different from an encyclopedia article. An encyclopedia article is free form text, and that's perfectly fine. But encyclopedia data is structured data inherently. You've got antonyms, you've got synonyms, you've got etymology. all of these components should be in tables in the database so that you can cross-reference and search them and things like that. So right now we're just doing it in a free-form wiki, although they are using templates to hopefully make it possible to convert to a better system later on. But the dictionary project does need more work. Free the curriculum.
I think that there will be a free, complete curriculum in every language, kindergarten through university level. If you think about this, this is really a much, much bigger task in the encyclopedia. A traditional encyclopedia is about yay big and all of the books that you would need to go from kindergarten up through the university level on all subjects is a much, much bigger library of books. So this is Wikibooks is underway, but there are a ton of other projects out there, some of which I think may be more successful than Wikibooks. There are
50:00–55:00
cases of physics professors who are banding together to just write textbook and they're using free license to do it, collaborating to build a textbook. And I think this is going to be a radical change for the textbook market. Right now, proprietary textbooks are controlled by a fairly small number of companies. There's sort of a star system. If you can get your economics 101 text adopted by a lot of places, then you're going to have a lot of sales. But although there's a handful of textbooks that are really, really, really popular,
there are tons and tons of professors who would be able to contribute to a collaborative effort to do it in a free and non-proprietary way. So I think this is going to happen fairly quickly. And the same thing applies not just to the university level, but think of school teachers who are not satisfied with the books that they're using. Any one school teacher would find it daunting to completely write a textbook from scratch, but you band together a couple hundred school teachers of sixth grade biology or something, and they can actually do something. So we have a question here, but I want to keep moving quickly so we can do a lot of questions at the end. At the K-12 level throughout most of the United States, teachers don't have discretion as to what textbooks they use.
Decisions are made by school boards or state textbook boards. Right. So that's actually a benefit and a hindrance to the K-12 effort because the benefit is that if you take a look at the California standards, for example, They're actually fairly detailed. The California standards for what a ninth grade history book should look like is a fairly long outline. And so that gives us a target that we can write to. Because one of the hard parts about writing a full textbook is figuring out what you should include or not include.
And that's a very difficult and complicated question if you approach it from scratch. For our community, it's much easier to say, here's the specification, let's meet the specification. Of course, since our work can be repurposed and reused, of course, people could then take that and mix and match and do whatever they want with it. But that's actually one thing. It provides us a target. One of the reasons Wikipedia works so well is if I say, encyclopedia article about the Eiffel Tower, everybody in the room has more or less the same idea of what that should look like. But if I say, you know, sixth grade history book, well, what should that be?
What should it cover? Those kinds of things. The hindrance, of course, is in adoption. That means if we want to get our textbooks adopted, we have to get them approved and on the list. And so we have to meet the specification and we also have to go through the approval process, which is costly and politically connected and all that sort of thing. So at the university level, it's much easier for professors to just pick and choose and experiment. Can I ask you to repeat questions and comments for the mic? Yes, yes, I will. So, okay.
Good. But Free the Music, the most amazing works in history are in the public domain, but not many public domain recordings exist. So you think about Mozart. Yes, the music itself is in the public domain, but as far as I know, there aren't really any public domain recordings. And then also proper scores, scores that meet modern standards for musicians, are also proprietary derivative works, an arrangement or something like this. And so I believe that we can have volunteer orchestras, student orchestras, community
orchestras. I doubt if we're going to see the really major orchestras, the ones who actually make money by selling CDs, are not going to release their work under a free license. But lots and lots of community orchestras, they release their work anyway. Maybe they sell a few CDs locally. It's really no big deal for them though to release that under a free license. So I think we're going to see this happen. if only those people are made aware of the possibility and people can bring all this together. Free the art. So these two paintings are from the National Portrait Gallery in London.
This is Shakespeare and Anne of Denmark. These paintings are 400 years old. And we got a really nice, well, a really nasty letter from them saying, we notice you have a number of images on your website which are of portraits in the collection of the National Portrait Gallery in London. Blah, blah, blah. Unauthorized reproduction of such content may be an infringement of our laws. Well, forget it. That's insane. Our response to that kind of threat is that the picture is still up on the website.
What I've taken to doing is I respond to these kinds of letters first with a couple of very boring paragraphs kind of explaining a lot of them, public domain, fair use, whatever the case may be. But then I put in a paragraph that says, you should be ashamed of yourselves. This is our shared cultural heritage. And I can often go to the museum's website and find their very noble statement of principles of how they're there to share knowledge and educate the public. And it's completely ridiculous. The behavior of museums in terms of trying to control the digital rights to the images in their collections are ridiculous.
One of the things that they've claimed is that, well, it's very, very expensive to make a proper digital copy of an image. and so da da da. That's complete nonsense. You give me the access and I'll have Wikipedians in there tomorrow. We'll do this for you. It's really no big deal. And so basically, our response to them is, you want to sue us? Bring it on because this is a case we would love to have. So far they just don't answer me. So we'll see what happens. Free the file
55:00–1:00:00
File formats. Proprietary file formats are worse than proprietary software because they live with no ability to switch at a later time. I think we're seeing a trend increasingly towards this idea. A lot of consumers are starting to realize it and as a lot of different kinds of software are available, people are starting to understand that every time you're forced to upgrade Microsoft Word because somebody sends you a document in the new format which you can't read, people start to get the idea I really shouldn't put all my data into something where it's being
controlled. A lot of companies are starting to recognize this as a huge problem. And there's considerable progress here and one of the most important parts of this progress is the continued rejection of software patents in the EU. This has been a really live and active debate within the EU about software patents. If you have a patent on a file format then it makes it really easy to lock it up as opposed to if your only method of protecting your file format is you don't explain this very crazy format, at least we have a chance of reverse engineering it.
So that's a really, really important thing. Software patents are a very bad threat to freedom. Freedom Maps. Stefan Magdalinski, a friend of mine in London who's done all kinds of cool stuff over there, when I told him about my idea for this list of 10 things, he said, what could be more public domain than basic information about location on the planet? There's a ton of GIS information, a lot of it's proprietary in the US. We're actually fortunate to have quite a bit that's available free because it comes from the US government. It's a public domain.
This is becoming increasingly important for open competition and mobile data services. One of the reasons that you're accessing information on your phone sucks so much is it's proprietary. They're trying to make money off of every little thing that you're doing. the ability to have this kind of free data, location data, will make it possible for people to program applications on phones where if you're standing at the Taj Mahal, you can push the button that says, what is this? And then it tells you, you're at the Taj Mahal.
Relax. So I think this is very important. Free the product identifiers. This one's kind of complicated to explain and we're a little short on time. I do want to get to some questions, so I'll just say it very quickly. The idea here, I was in Copenhagen and I met Ulla Maria Mutanen, who is a PhD student in fashion. And I thought, oh, this is interesting. What do I have in common with someone who's getting a PhD in fashion? But it turns out that her PhD work is into the phenomenon of people crafting and selling crafts online.
And there's a whole long tail of people who are doing all kinds of interesting work. But when they go to sell their products into the global market, all they can really do is use their own little website and things like this. What they really need is a method of getting a product identifier so that their work can be identified. And then, therefore, you can build all kinds of e-commerce platforms around this so people can do aggregation of that kind of information. Right now, one of the things that's going on is, for example, Amazon has their Amazon
identifier number. They also accept ISBNs, which is the standard book number. So my concept is that we should have what I would call long tail identification numbers. So this is a free system where people can get these numbers, which could be then part of a standard system. And there's a lot of people working on this all around the world to develop a standard for this kind of identifier. Free the search engine. Transparency in finding things is critically important. Free search technology exists but needs work.
And since this is my list of things that I think will be free, I think that it's possible to have a commercial-non-commercial hybrid that could be transparent and self-sustaining. So the idea here would be, the economic model of such a thing, would be something like Google. I probably shouldn't tell this in Silicon Valley because somebody will rush out and raise money from a VC and not tell me about it. but here I am, what am I going to do? The commercial, non-commercial hybrid would be you use the model for funding it like Google, so you use ad revenue,
but unlike Google and all the other major search engines, you fully publish all your algorithms, and you fully publish how you're doing everything. Now, you might imagine that, well, that's our core proprietary thing. We have all the best stuff, and we don't want to give it away, because then why would anybody come to us? But depending on the type of license that you use, if you require attribution and things like this, it could work exactly the same way as Wikipedia works. That, yes, although Wikipedia gives everything away, it's a hugely popular website. And so even a search engine that gives everything away could be a hugely popular website.
1:00:00–1:05:00
Listen from 1:00:19 (MP3 audio)
You don't make money from selling the technology. You make money from consumer trust in the technology. So I think that's something that is likely to happen. And free the community. So Wikipedia demonstrates the power of a free community. I believe that consumers of web forums, wiki services, mailing lists, all kinds of things on the internet, should demand a free license, a license to their work that enables them to control the destiny of their own community. Otherwise the company controls the community. I did some consulting at the BBC. The BBC is a wonderful organization.
Listen from 1:00:51 (MP3 audio)
They really take seriously their mission to the public. They really want to work to make their website more interactive. They want to let the community come in and kind of control what's going on. All those kinds of things. But if you read the terms of service of the BBC website, every single thing you submit, they have a completely abusive, but very common, terms of use that say basically they own everything you do on the site forever. Period. Instead, what people should do if you've got a community is take some care for what happens if the company starts being weird on us.
Listen from 1:01:25 (MP3 audio)
Do we have a way out? Do we have a way? Are we allowed to take our own work as a group without getting permission from every last single person to do it? Is there a practical way that we could actually fork and leave? And I think we're starting to see an increased awareness of this kind of thing from a lot of people on the internet that as you see different services go out of business or the company changes their mind about something and people realize, you know, we don't really control the fate of our community unless we control our work. And so the question would be, are you a serf living on your master's estate or are you free to move?
Listen from 1:01:56 (MP3 audio)
That's the fundamental question I think people are going to ask themselves. So that's essentially the end of my talk, and I'm very happy to have questions now. Yeah? So I know of at least two judicial opinions that cite Wikipedia, and at least two legal briefs, things written by lawyers, two forms. And none of these made use of the feature of Wikipedia where you can actually link to a specific instance of a fake. I'm wondering if, do you think that there needs to be better education about this, about how to use that kind of feature?
Listen from 1:02:30 (MP3 audio)
Right. And in general, about Wikipedia norms, I found it really kind of hard to sort of get a dispute about NPOV and things like that. Right. So the question is, there have been cases of fairly important things like courts who are citing Wikipedia, and they're not citing particular revisions in the revision history, so it really doesn't make sense. what they cited could easily be changed and could say the opposite at some point in the future. And then in general, so should there be more education about how to use Wikipedia?
Listen from 1:03:01 (MP3 audio)
And I'd say yes. I think people do need to be more educated about it. The whole question of reputation, authority, citing Wikipedia, all those kinds of things are very deeply interesting questions. The Wikipedia community, the core community, is deeply committed to high quality. Yet our work is not high quality in every respect. It's quite good in many ways, but we know it's a work in progress. And so it is a little uncomfortable.
Listen from 1:03:32 (MP3 audio)
About once a week now, and it's increasing frequency, I get an email from some university student somewhere who says, I quoted Wikipedia in my paper and I got an F, can you help me? And I just answer, you know, you're in college. I don't know why you're citing any encyclopedia. I hope they don't let you cite Britannica. That's not the purpose of an encyclopedia. The purpose of an encyclopedia is to give you a broad background, some context. So you're assigned some reading for class, and maybe you're having some trouble with some of the background information.
Listen from 1:04:04 (MP3 audio)
It references a place or a person, and you don't really know who that is or what that is, and you're having trouble with the context. Then you go look it up in Wikipedia, and you get your basic context. Or, in a different kind of realm, if all you need is to kind of know something, then you can go kind of read it in Wikipedia, and that's fine. That's the way I think most people use Wikipedia. It's just kind of an intelligence booster. You just sort of need to know something, and you go read a couple paragraphs, and you're up to speed. The idea of using any encyclopedia, much less one that can be changed every instant, as some kind of a stable academic reference, this doesn't make any sense to me.
Listen from 1:04:37 (MP3 audio)
So I think people will come to understand that, but I hope they come to understand it in a positive way, that this isn't saying anything bad about Wikipedia. This is saying particular kinds of sources have their particular kind of role. You said one thing about putting images on a separate server in part, I think, to deal with local copyright issues. Ah, okay, right. Yeah, I guess now that I think back on what I said,
1:05:00–1:10:00
Listen from 1:05:09 (MP3 audio)
I didn't come to the punchline of that and clarify that. Because we have to deal with multiple jurisdictions, at Wikimedia Commons, we don't allow fair use, and we follow sort of a superset of all laws that we can. So the idea is that all of the images on Wikimedia Commons should be usable in every language of the world. There are, of course, exceptions, but those exceptions have to do with censorship, not with copyright. So fair use isn't a reasonable reason to put images there. It has to be public domain or...
Listen from 1:05:44 (MP3 audio)
Right. Public domain or freely licensed. It has to be released correctly. In the English Wikipedia, we do use a fair number of fair use images. In my opinion, we are moving in the direction of using fewer and fewer and fewer of those. There are real problems with fair use images. One, the legal doctrine of fair use applies very, very clearly to what we're doing. We're non-profit. We're educational. We're taking small images. If you look at the four-factor test, our uses of fair use are quite clean.
Listen from 1:06:17 (MP3 audio)
On the other hand, we want people to be able to reuse our content. So we want people to do it commercially. That's one of the things we want. Also, we want people to be able to slice out parts of our work and use it for whatever. It could be all kinds of different purposes. Another big problem with fair use, a friend of mine is a baseball fanatic. He goes to lots and lots of baseball games and he takes pictures of baseball players. He doesn't have a press pass so he has to stand up in the stands. He has a nice camera and he's pretty good at it. But at the same time it's a little disheartening to him to go to all this work.
Listen from 1:06:49 (MP3 audio)
Then he goes into Wikipedia and he sees in the article a photo of somebody nicked off the web somewhere. Or a still frame from a video or something. By having excessive reliance on fair use, we discourage people from actually coming up with free alternatives. So I think we're in the process of really rethinking how we use fair use. And then additionally, although we generally follow the US law for fair use in the English Wikipedia, the US isn't the only country that speaks English. So we really want our work to be reusable in the UK, for example.
Listen from 1:07:23 (MP3 audio)
Or anywhere in the world, people might find the English useful. So, yeah. The tools that you mentioned for monitoring that people develop and for investigating due to this behavior, are all those tools that are source code available publicly? And is the same, the determinations on who has access to actually run those tools goes the same sorts of public governments?
Listen from 1:07:56 (MP3 audio)
Right. Yes. Yes, but all that should clarify. So any kind of tool that might be developed by any member of the general public, they can put it under whatever license they want. I don't know of any examples of people developing a proprietary tool than expecting the Wikipedia community to use it. That just would be very strange for us. So these are all free software programmers and things who are just releasing their stuff under GPL. In terms of who can use these, these aren't tools that have any kind of special access to the website per se. The one that I was actually mentioning is CryptoDirk's Vandal Fighter.
Listen from 1:08:30 (MP3 audio)
It works by reading the live feed. It's a very public IRC channel. Anyone can go to the IRC channel and listen in. It's just a little inconvenient because it's quite noisy. That tool doesn't have any special access. But there was another one, the Check User tool. The check user tool, yeah, that's freely licensed, but it's on our website. It actually accesses the database and the log files and things like that. And so, yeah, it's freely licensed so that other people who have a large wiki project could use our tools if they're using our software.
Listen from 1:09:02 (MP3 audio)
But then the control to access that is through the community process. Yeah, back in the back. Can you... Yes? Can you tell us a little bit about fundraising or financial? I mean, although this is a non-profit organization, but I'm sure you guys need a little more amount of money. Yeah, so we occasionally, the question is, can you tell a little bit more about fundraising and the finances of the project and things like that?
Listen from 1:09:34 (MP3 audio)
Occasionally we'll have a fun drive on the site where we say to people, you know, we're trying to give away free encyclopedias, so please help us. My feeling is that unlike almost any other non-profit, because we are a community, we're unbelievably tasteful about asking for money. I always feel we could be a little bit more panic-stricken because that seems to be the tradition for non-profits. But, you know, donate now or we're closing Wikipedia. So I think that would bankrupt some of our really addicted users if we said that. So we don't want to abuse them.
1:10:00–1:15:00
Listen from 1:10:06 (MP3 audio)
But we've always had a really good success when we need to raise money from the public. We had, you know, several months ago we had a fun drive. We had set ourselves three weeks to reach our goal, and we had to take the banners down off the site after two weeks because we reached the goal. We just recently raised, we were trying to raise $200,000. We actually raised closer to $250,000 in just a few weeks' time. We also are getting, we have very favorable response from grant-making institutions,
Listen from 1:10:40 (MP3 audio)
although the problem that we've found so far there is that your large philanthropic organizations are very bureaucratic, and they don't really know how to deal with us. I met actually the head of the Dutch Minister of Culture. She oversees a 30 million euro budget. She's a huge fan of Wikipedia and she says we would love to help you. But what they end up funding is they fund the Rijksmuseum because it's really easy to
Listen from 1:11:13 (MP3 audio)
understand. They fund the Anne Frank House. It's simple. There's things to pay for and they just pay for that. When you look at us and it's just, well back then when I had the conversation with her, like this goofy guy with an office at home and a bunch of servers and a whole bunch of crazy people on the internet. It's kind of hard to see how can they fund that or whatever. As we're growing, we're becoming more mature as an organization. So we're beginning to be able to actually make proper grant applications. And we believe we'll have very good success there. And then additionally, we're getting very good response from potential corporate sponsors.
Listen from 1:11:46 (MP3 audio)
So Yahoo, for example, was very happy to give us some service for our facility in South Korea. They do it for all the reasons that any company might help a non-profit. It's good PR for them. The other thing I think, one of the things I say is that the major search engines, for example, so Yahoo and Google, MSN, I doubt if they'll ever get any money from Microsoft, but we'll see. Their whole business model fundamentally depends on the internet not sucking.
Listen from 1:12:18 (MP3 audio)
And one of the things that we do is make the internet not suck. So they really should throw us a bone now and then. We've had pretty good response from people who say, yeah, we love Wikipedia, it doesn't cost that much, so we can give you some money to do it. So I'm very optimistic about the future of fundraising and financing. One of the principles that we're using there is we really strongly value our independence. So if we get an offer from a big company, they'll host everything, they'll do everything for us, we'll turn down that offer. Because it makes us too dependent on one organization, which isn't really good for us in the long run.
Listen from 1:12:54 (MP3 audio)
But accepting help from lots of different kinds of people and parties is a really good idea. Okay, here. I'm wondering about what issues you're having with scaling. You've already identified that there's a relatively small community of users that are making most of the edits. Is that growing and you're starting to bring into issues with scaling up the system? It is, yeah. We have had some problems with scaling. One of the main problems we've had with scaling is that I am personally exhausted.
Listen from 1:13:25 (MP3 audio)
So that's one of the reasons that we're really trying to push as much things off of me as possible. What I say about this is that one thing that we know in general is that communities inherently scale. We know this because ordinary traditional human communities scale. You can live in a small village or a small town or a small city or a big city, things scale. So when you don't have a centrally planned idea in mind, you have interactions between lots of people in little neighborhoods and things, it can actually scale.
Listen from 1:13:57 (MP3 audio)
but as with the growth from a small village to a large city you get new kinds of problems, you get new kinds of difficulties because you've got more and more people around maybe the old way of doing things doesn't scale anymore the law maybe has to become a little less informal and a little more formal so yeah, whenever people ask me how do you see the future of the community in terms of what happens when instead of saying you've got X number of users when you've
Listen from 1:14:30 (MP3 audio)
got ten times that number working. I would say well for the English Wikipedia I really don't know. All I know is we've got a lot of really smart, really thoughtful people who are thinking and discussing this all the time and we're not going to let it slip away from us. But it's very hard to predict how we're going to deal with various challenges. One of our rules that we've always had is we don't solve problems before they actually happen. Because that kind of excessive a priori thinking is really what kills a lot of this kind of thing. One of my rants is about designers of social software solving problems that don't exist.
1:15:00–1:17:33
Listen from 1:15:09 (MP3 audio)
The example I give is suppose you're designing a restaurant. and in this restaurant you're going to be serving steak and so since you're gonna have steak you're gonna have steak knives so if you have knives people might stab each other so therefore you need cages around all the tables right that's a crazy chain of thought because think about what does that do to your civil society what does it do your community it's obviously crazy that's the way most social software is designed that you think of all the bad things the user might do and you make sure you've got a very complex permissions model to make
Listen from 1:15:41 (MP3 audio)
sure they don't do things. The wiki approach is to go the exact opposite and say, you know, people probably won't stab each other, and if they do, well, okay, you know. There's a huge number of benefits. The same reason we don't say, well, in order to make sure we're all safe from being stabbed, we're going to put cages in all the restaurants. We say, well, take that risk, because the joy of sitting in a, you know, the simple human joy of sitting in a restaurant next to someone who's eating a steak and they're not stabbing you with it is nice. We don't give it much thought because it's so normal. One of the things I like to say is that I think
Listen from 1:16:15 (MP3 audio)
partly because of the the personality types who become programmers, I don't know what it is exactly, but a lot of programmers seem to me to think that the whole point of social software is to replace the social with the software, which is not really what you want to do, right? Social software should exist to empower us to be human, to interact in all the normal ways that humans do. And so that's one of our design principles that I think will help us scale better in the future.
Listen from 1:16:48 (MP3 audio)
But it's also one of the reasons I don't have any answers yet. Because we don't know what's going to happen, and so we can't solve it until it actually happens. So I've got to run, because I've got to go to the city for a Creative Commons event. And I'm the guest of honor, and I'm going to be like, late. So I'm sorry I can't stay longer. We did record this talk. You can talk to me or you can subscribe to
Listen from 1:17:20 (MP3 audio)
I'm a founder of the I just wanted to meet you. I hope you see a bit later. I didn't want to meet you while I had you here. Thanks a lot for all your work.
Generation and audio provenance
- Generated
- 2026-09-09T22:49:16.215238+00:00
- Speech-recognition engine
- mlx-whisper 0.4.3
- Model
- mlx-community/whisper-large-v3-turbo
- Model revision
a4aaeec0636e6fef84abdcbe3544cb2bf7e9f6fb- Model weights SHA-256
951ed3fc1203e6a62467abb2144a96ce7eafca8fa77e3704fdb8635ff3e7f8a6- Original audio file
- wales_sims_03-Nov-2005.mp3 (38103531 bytes)
- Original audio SHA-256
e8c9f1e1ce8b42f44519b2800d835db70cce29a84dd0cd568b0abf32ae7e3459- Completion receipt SHA-256
1165fe3c3394d58dcb1430dea86a597022b18dcdb86b9b042f9892103bd72ee9- Normalized generated text SHA-256
bb244a2d99f5ee19e1143f0190a3e56d139f7c0db8d6d30913c016b673876305
Audio stayed local during generation. This provenance identifies the source and process; it does not certify transcription accuracy.
Back to the top of this transcript · Return to the original talk post