logo
EverydayChaos
Everyday Chaos
Too Big to Know
Too Big to Know
Cluetrain 10th Anniversary edition
Cluetrain 10th Anniversary
Everything Is Miscellaneous
Everything Is Miscellaneous
Small Pieces cover
Small Pieces Loosely Joined
Cluetrain cover
Cluetrain Manifesto
My face
Speaker info
Who am I? (Blog Disclosure Form) Copy this link as RSS address Atom Feed

June 22, 2011

Ethanz to head the MIT Civic Media Center

Ethan Zuckerman has been named the new director of MIT’s Center for Civic Media.

This is fantastic news for the Center. There is no one imaginably better for this position than Ethan. Plus, Ethan and Joi Ito [twitter:joi] (new head of MIT’s Media Lab) will be working together, which promises a type of quantum energy not seen since the Big Bang.

Ethan of course will have less time to spend at the Berkman Center, where he is an irreplaceable source of heart and brains. But, he’s actually going to be in Cambridge considerably more than he has been (he commutes from western Mass.), and the two centers are already discussing deeper, richer collaborations.

I am privileged to count Ethan as a close friend, and I couldn’t be happier for him. Plus, the collective, collaborative energy emanating from Cambridge is about to multiply. Woohoo!

Tweet
Follow me

Categories: berkman Tagged with: berkman • c4 • ethan zuckerman Date: June 22nd, 2011 dw

Be the first to comment »

June 16, 2011

First federal CIO is coming to Berkman!

Vivek Kundra, of whom I am a fanboy, is leaving his position as our first federal CIO. That’s too bad for America. I think he has done an outstanding job.

But he’s coming to the Berkman Center as well as the Kennedy Center at Harvard as a Fellow, which is great news for us.

Tweet
Follow me

Categories: berkman, open access Tagged with: berkman • vivek kundra Date: June 16th, 2011 dw

Be the first to comment »

June 11, 2011

Berkman Buzz

This week’s Berkman Buzz:

  • Wendy Seltzer [twitter:wseltzer] explores privacy in public: link

  • Ethan Zuckerman [twitter:ethanz] navigates privacy walls and thresholds: link

  • Dan Gillmor [twitter:dangillmor] reviews the FCC report on the future of media: link

  • The OpenNet Initiative evaluates what a new treaty on gambling could mean for Internet filtering in Germany: link

  • Weekly Global Voices [twitter:globalvoices] : “Syria: True Identity of Arrested Blogger Questioned”: link

Tweet
Follow me

Categories: misc Tagged with: berkman Date: June 11th, 2011 dw

Be the first to comment »

May 28, 2011

Weekly Berkman Buzz

This week’s Berkman Buzz:

  • Ethan Zuckerman [twitter:ethanz] explores the lessons to be learned from an Azeri journalist’s release from jail:
    link

  • Harry Lewis compares $100,000 to a college education:
    link

  • The Citizen Media Law Project [twitter:citmedialaw] explains a British court order against Twitter:
    link

  • Howard Rheingold [twitter:hrheingold] covers the impact of digital media on youth civic engagement for DMLcentral:
    link

  • Chris Soghoian [twitter:csoghoian] reviews the DOJ’s reinterpretation of the Patriot Act:
    link

  • Weekly Global Voices [twitter:globalvoices] : “Russia: Attack Survivor Journalist Oleg Kashin on Internet Freedom”:
    link

Tweet
Follow me

Categories: berkman Tagged with: berkman Date: May 28th, 2011 dw

1 Comment »

May 13, 2011

Berkman Buzz

This week’s Berkman Buzz:

  • Wendy Seltzer [twitter:wseltzer] inspects son-of-COICA:
    link

  • OpenNet Initiative reports on the Syrian Electronic Army, and Facebook:
    link

  • Media Cloud investigates Russian blogs, media and agenda-setting:
    link

  • Ethan Zuckerman [twitter:ethanz] keynotes CHI 2011 — parataxis, cities, serendipity, design:
    link

  • David Weinberger discusses e-books and much more with James Bridle:
    link

  • Citizen Media Law Project [twitter:citmedialaw] introduces us to the OpenCourt project:
    link

  • Weekly Global Voices [twitter:globalvoices] : “Uganda: Museveni’s Swearing in Overshadowed by Rival’s Return”
    link

Tweet
Follow me

Categories: berkman Tagged with: berkman Date: May 13th, 2011 dw

Be the first to comment »

May 10, 2011

[berkman] Culturomics: Quantitatve analysis of culture using millions of digitized books

Erez Lieberman Aiden and Jean-Baptiste Michel (both of Harvard, currently visiting faculty at Google) are giving a Berkman lunchtime talk about “culturomics“: the quantitative analysis of culture, in this case using the Google Books corpus of text.

NOTE: Live-blogging. Getting things wrong. Missing points. Omitting key information. Introducing artificial choppiness. Over-emphasizing small matters. Paraphrasing badly. Not running a spellpchecker. Mangling other people’s ideas and words. You are warned, people.

The traditional library behavior is to read a few books very carefully, they say. That’s fine, but you’ll never get through the library way. Or you could read all the books, very, very not carefully. That’s what they’re doing, with interesting results. For example, it seems that irregular verbs become regular over time. E.g., “shrank” will become “shrinked.” They can track these changes. They followed 177 irregular verbs, and found that 98 are still irregular. They built a table, looking at how rare the words are. “Regularization follows a simple trend: If a verb is 100 times less frequent, it regularizes 10 times as fast.” Plus you can make nice pictures of it:


Usage is indicated by font size, so that it’s harder for the more used words to get through to the regularized side.


The Google Books corpus of digitized text provides a practical way to be awesome. Erez and Jean-Baptiste got permission from Google to trawl through that corpus. (It is not public because of the fear of copyright lawsuits.) They produced the n-gram browser. They constructed a table of phrases, 2B lines long.


129M books have been published. 18M have been scanned. They’ve analysed 5M of them, creating a table with 2 billions rows. (In some cases, the metadata wasn’t good enough. In others, the scan wasn’t good enough.)

They show some examples of the evolution of phrases, e.g. thrived vs. throve. As a control, they looked at 43 Heads of State and found that the year they took power usage of “head of state” zoomed (which confirmed that the n-gram tool was working).


They like irregular verbs in part because they work out well with the ngram viewer, and because there was an existing question about the correlation of irregular and high-frequency verbs. (It’d be harder to track the use of, say, tables. [Too bad! I’d be interested in that as a way of watching the development of the concept of information.]) Also, irregular verbs manifest a rule.


They talk about chode’s change to chided in just 200 yrs. The US is the leading exporter of irregular verbs: burnt and learnt have become regular faster than others, leading the British’s usage.


They also measure some vague ideas. For example, no one talked about 1950 until the late 1940s, and it really spiked in 1950. We talked about 1950 a lot more than we did, say, 1910. The fall-off rate indicates that “we lose interest in the past faster and faster in each passing year.” They can also measure how quickly inventions enter culture; that’s speeding up over time.


“How to get famous?” They looked at the 50 most famous people born in 1871, including Orville Wright, Ernest Rutherford, Marcel Proust. As soon as these names passed the initial threshhold (getting mentioned in the corpus as frequently as the least-used words in the dictionary) their mentions rise quickly, and then slowly goes down. The class of 1871 got famous at age 34; their fame doubled every four years; they peaked at 73, and then mentions go down. The class of 1921’s rise was faster, and they became famous before they became 30. If you want to become famous fast, you should become an actor (because they become famous in the mid to late 20s), or wait until your mid 30s and become a writer. Writers don’t peak as quickly. The best way to become famous is to become a politician, although have to wait until you’re 50+. You should not become an artist, physicist, chemist or mathematicians.


They show the frequency charts for Marc Chagall, US vs. German. His German fame dipped to nothing during the Nazi regime who suppressed him because he was a Jew. Likewise with Jesse Owens. Likewise with Russian and Chinese dissidents. Likewise for the Hollywood Ten during the Red Scare of the 1950s. [All of this of course equates fame with mentions in books.] They show how Elia Kazan and Albert Maltz’s fame took different paths after Kazan testified to a House committee investigating “Reds” and Maltz did not.


They took the Nazi blacklists (people whose works should be pulled out of libraries, etc.) and watched how they affected the mentions of people on them. Of course they went down during the Nazi years. But the names of Nazis went up 500%. (Philosophy and religion was suppressed 76%, the most of all.)


This led Erez and Jean-Baptiste to think that they ought to be able to detect suppression without knowing about it beforehand. E.g., Henri Matisse was suppressed during WWII.


They posted theirngrams viewer for public access. From the viewer you can see the actual scanned text. “This is the front end for a digital library.” They’re working with the Harvard Library [not our group!] on this. In the first day, over a million queries were run against it. They are giving “ngrammies” for the best queries: best vs. beft (due to a character recognition error); fortnight; think outside the box vs. incentivize vs. strategize; argh vs aargh vs argh vs aaaargh. [They quickly go through some other fun word analyses, but I can’t keep up.]


“Cultoromics is the application of high throughput data collection and analysis to the study of culture.” Books are just the start. As more gets digitized, there will be more we can do. “We don’t have to wait for the copyright laws to change before we can use them.”


Q: Can you predict culture?
A: You should be able to make some sorts of predictions, but you have to be careful.


Q: Any examples of historians getting something wrong? [I think I missed the import of this]
A: Not much.


Q: Can you test the prediction ability with the presidential campaigns starting up.
A: Interesting.


Q: How about voice data? Music?
A: We’ve thought about it. It’d be a problem for copyright: if you transcribe a score, you have a copyright on it. This loads up the field with claimants. Also, it’s harder to detect single-note errors than single-letter errors.


Q: Do you have metadata to differentiate fiction from nonfiction, and genres?
A: Google has this metadata, but it comes from many providers and is full of conflicts. The ngram corpus is unclean. But the Harvard metadata is clean and we’re working with them.


Q: What are the IP implications?
A: There are many books Google cannot make available except through the ngram viewer. This gives digitizers a reason to digitize works they might otherwise leave alone.


Q: In China people use code words to talk about banned topics. This suppresses trending.
A: And that takes away some of the incentive to talk about it. It cuts off the feedback loop.


Q: [me] Is the corpus marked up with structural info that you can analyze against, e.g., subheadings, captions, tables, quotations?
A: We could but it’s a very hard problem. [Apparently the corpus is not marked up with this data already.]

Q: Might you be able to go from words to metatags: if you have cairo, sphinx, and egypt, you can induce “egypt.” This could have an effect on censorship since you can talk about someone without using her/his name.
A: The suppression of names may not be the complete suppression of mentions, yes. And, yes, that’s an important direction for us.

Tweet
Follow me

Categories: berkman, copyright, too big to know Tagged with: 2b2k • berkman • google • irregular verbs • library Date: May 10th, 2011 dw

2 Comments »

May 6, 2011

News is a wave

By coincidence, here are two related posts.

Gilad and Devin at Social Flow track the enormous kinetic energy of a single twitterer who figured out shortly before President Obama’s announcement that Osama Bin Laden had been killed. But my way of putting this — kinetic energy — is entirely wrong, since it was the energy stored within the Net that propelled that single tweet, from a person with about a thousand followers, across the webiverse. And the energy stored within the Net is actually the power of interest, the power of what we care about.

Meanwhile, The Berkman Center today announced the public availability of Media Cloud, a project Ethan Zuckerman and Hal Roberts led. Ethan explains it in a blog post that begins:

Today, the Berkman Center is relaunching Media Cloud, a platform designed to let scholars, journalists and anyone interested in the world of media ask and answer quantitative questions about media attention. For more than a year, we’ve been collecting roughly 50,000 English-language stories a day from 17,000 media sources, including major mainstream media outlets, left and right-leaning American political blogs, as well as from 1000 popular general interest blogs. (For much more about what Media Cloud does and how it does it, please see this post on the system from our lead architect, Hal Roberts.)

We’ve used what we’ve discovered from this data to analyze the differences in coverage of international crises in professional and citizen media and to study the rapid shifts in media attention that have accompanied the flood of breaking news that’s characterized early 2011. In the next weeks, we’ll be publishing some new research that uses Media Cloud to help us understand the structure of professional and citizen media in Russia and in Egypt.

Now Media Cloud is going to be a very useful tool. And it was not trivial to build. Congratulations to the team. And thank you.

Tweet
Follow me

Categories: berkman, media Tagged with: berkman • media • news Date: May 6th, 2011 dw

1 Comment »

April 29, 2011

Berkman Buzz

This week’s Berkman Buzz:

  • danah boyd gets kicked off tumblr by a company and writes about getting her identity back:
    link

  • Doc Searls chronicles the recent public debate about personal data:
    link

  • Dan Gillmor does not support the blogger lawsuit against Huffington Post:
    link

  • Ethan Zuckerman explores the roles facts and values play in polarization:
    link

  • Stop Badware is developing best practices for malware reporting:
    link

  • Weekly Global Voices: “Rwanda: Ask Rwandan President Questions on YouTube”:
    link

Tweet
Follow me

Categories: berkman Tagged with: berkman Date: April 29th, 2011 dw

3 Comments »

April 22, 2011

Berkman Buzz

This week’s Berkman Buzz:

  • Ethan Zuckerman [twitter:ethanz] documents Internet filtering at the National Science Foundation: link

  • Radio Berkman talks to Steven Levy about the Googleplex: link

  • The OpenNet Initiative covers the Ugandan government’s new Internet filtering attempts: link

  • The Citizen Media Law Project [twitter:citmedialaw] reviews recent Righthaven copyright cases: link

  • Weekly Global Voices [twitter:globalvoices] : “Chile: Nurse Expedites Organ Transport Using Twitter”: link

  • Tweet
    Follow me

    Categories: berkman Tagged with: berkman Date: April 22nd, 2011 dw

    Be the first to comment »

April 19, 2011

[berkman] Protocol.by

Greg Elliott and Hugo van Vuuren are giving a Berkman talk on “The Communication Crises and the Evolution of Personal and Cultural Protocols.” They are launching a new tool this week: Protocol.by. (Ethan Zuckerman has posted his live blogging of this talk.)

NOTE: Live-blogging. Getting things wrong. Missing points. Omitting key information. Introducing artificial choppiness. Over-emphasizing small matters. Paraphrasing badly. Not running a spellpchecker. Mangling other people’s ideas and words. You are warned, people.


They begin with a video that talks about the number of channels and messages in which we’re drowning. This is the communication crisis Greg and Hugo are addressing. They are interested in how we deal with the guilt of (Tina Roth Eisenberg) of not being able to keep up. We have various tools, such as email bankruptcy. They point to an XKCD Map of Online Communities that, among other things, reminds us that the Net is dwarfed by other forms of communication. A NY Times article (March 18, 2011) is about our culture’s movement away from telephone calls, even though you get more metadata; in many instances, a quick text is more appropriate.


The Internet is a Rorschach test, they say. We all play a puzzling game with our email, trying to filter it without missing anything and without hurting anyone’s feelings. E.g., danah boyd famously takes email sabbaticals, during which her auto-responder tells you that she will never read your msg. Other people (including Tim Berners-Lee) have detailed instructions about the netiquette for contacting him.


The site five.sentenc.es was influential on Greg and Hugo. It provides a link for your sig that announces that all your email responses will be five sentences or less. We used to have posters that instruct children in good manners. They’re not proposing that, of course. But, announcing norms shapes behavior.


They show their site: Protocol.by. (Sign up here: http://protocol.by/newUser, and give them time to hand-approve you.) Once you sign up, you create a profile that tells people your preferences in being contacted: Which channels, in which priority, and expectations. E.g., Use email; it may take me a while to get back to you and you don’t need to wrap it in social niceties; if necessary, call my phone, but don’t leave a message. (Here’s my profile.) This is even more useful, they say, if plugged into a community as a group protocol.


They are gathering data for research into how people rank their channels. (Anonymous, of course.) (Greg points to the data at the okCupid dating site.)


Q: What’s your business model?
A: This is a side project. We’re in it for the research.


Q: There’s a risk in making these rules too explicit. E.g., it says you respond in 24 hours, but you never want to respond to some particular person and they then get offended.
A: We encourage users to leave in as much ambiguity as they can. It’s up to you the user to define it.


Q: So much of the preferred channel is based on who the person is: If you’re my babysitter I want you to call, but if you’re my grad student, use email. How are those directions indicated?
A: We can imagine the site presenting different protocols depending on who you are: Are you a stranger, are you a friend? For now, Protocol is aimed at strangers since your friends probably already know how to reach you.


Q: How many users do you need for research purposes, and how are you going to get them?
A: We have 500-600 already. A big sample would be thousands. The next step is the location setting, and embedding into other services. We also want to reach people who are already using these sorts of rules.


Q: I love that you’re providing a tech solution, but are talking about the human problems. We are now past the era of flaming. Has your data shown if these protocols help prevent people from getting offended?
A: We don’t have the data yet.


Q: I’m a huge fan because it brings peace of mind. Each new channel fragments our identity. I love that Protocol centralizes our communicational identity no matter how our technology changes. Your suggesting that our communicational identity is our social identity. How is our identity crafted by our communication tools?
A: Yes, our identities are shaped by our tools. But I don’t know that Protocol is going to shape our identities or represent it. Some users do have very specific rules, which they use as a signal that they are very busy.


Q: You have a distinct individualistic bias. You think we’re going to pick our own tools and ways of communicating. You’re young and tech savvy. But I deal with the press, and they’re going to call my phone no matter what I say, because they have more power than I do. I wonder if asserting these protocols is a transitional moment. Maybe we’ll centralize on a new socially acceptable set of protocols, or are we going to fragment?
A: Communication media don’t generally replace predecessors. We’re not going to a singularity of communication preferences, but it will boil down to a smaller set. E.g., Rapportive (gmail extension) fetches info about the sender of any email msgs — their twitter account, etc.


Q: You said your motivation is relieve guilt. This seems like a geeky way to deal with the social anxieties that geeks tend to have.
A: In the future, we need systems to offer the protocols without you having to seek them out.


Q: There has to a brand of new psychologists dealing with these issues: When something is ambiguous, does it mean someone hates me, etc.?
A: Yes. Interesting.


Q: When you don’t know someone, I’d probably google them and find their primary-facing piece of info. How do you get Protocol to become that piece of info, especially when you’re talking about different ages, communities?
A: Embedding, for one thing.

Q: We have collapsed boundaries between channels. I want some people to self-declare what subjects they’re interested in. I’d rather tell people what I’m interested in rather than have them mine it and guess.
A: The word “reputation” hasn’t gone up yet. It used to matter more when we lived in small communities. Now we can invent ourselves many times. As the Net goes into its next phase, reputation and data will matter a great deal.

Q: Embedding is a nice idea. Get some sites to embed a cute logo. Second, you’re increasing the velocity, but velocity is the problem. There’s no barrier on the sender’s side to communication. Is anyone talking about putting actual costs on email. It should cost people to email me. That would slow the velocity.
A: We see Protocol as being the barrier eventually. The problem with money is that it discriminates invidiously. Also, it’d be nice if I could ping Protocol to see if my friend is available to talk, and it knows enough about our relationship and his circumstances. There’s no cost, but your msg may not get through.

Q: It used to be easy. Now you may not want to let people know why the time zone you’re in. You might want to have an abstraction layer that knows the zone you’re in and what the preferred order is in various zones.
A: There are many variables. The issue is that you get into complex, power-user

Q: [me] There will be an increasing need for metadata because the community of possible communicators has increased, with greatly differing local norms. So, how about creating a little marker that lists your preferred channels order, the way Creative Commons lets you easily represent your license preferences. Then let institutions encourage their users to put the marker on their web pages, etc.
A: We’re thinking that we’ll have three markers with varying degress of info.
Q: Well, one marker is easier to market than 3.

How about having two profiles, so you can give your friends the key?
A: You get into DRM issues. We’d rather encourage people to use Protocol for non-friends.


Q: How will you keep it up to date with new channels/services?
A: We hope that Protocol will be a part of the channel tools you use.


Q: Privacy used to be the right to be left alone. Now it’s about how much info is out there and how you can control it. You seem to adopt somewhat of a technological determinist standard; tech determines our social norms. Is that what you’re driving at?
A: We don’t think tech is the defining determinant of our lives, but it is there. We either let it run wild or set some rules.

Tweet
Follow me

Categories: berkman Tagged with: berkman Date: April 19th, 2011 dw

1 Comment »

« Previous Page | Next Page »


Creative Commons License
This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 3.0 Unported License.
TL;DR: Share this post freely, but attribute it to me (name (David Weinberger) and link to it), and don't use it commercially without my permission.

Joho the Blog uses WordPress blogging software.
Thank you, WordPress!