2009-01-17
2009-01-04
2009-01-01
Randy Pausch Time Management 0:24.45 Does anyone here have an e-mail inbox sorted by importance? This is my PhD topic. Defining what importance means and applying it to the items in the universal inbox.
Designing a More Effective Inbox has some statistics but I don't think they have the solution. They also use utilize or utilized five times which makes me suspicious of their claims. They also references but no in-text citations and claim to have patented a process or algorithm. Paul Graham has an interesting essay on software patents.
Franklin, M, Halevy, A & Maier, D (2005) From Databases to Dataspaces: A New Abstraction for Information Management might have relevance to online business group behaviour. I know there is a lot of work on organisational behaviour that I need to reference but mostly stay away from.
2008-12-01
2008-09-21
2008-09-07
2008-07-31
John P.A. Ioannidis Why Most Published Research Findings Are False. PLoS Medicine August, 2005;2;8:696-700
2008-07-29
2008-06-18
2008-06-11
Interesting presence ideas
2008-06-06
2008-05-28
2008-05-27
2008-05-18
Jean-Yves Le Boudec Artificial Immune System For Collaborative Spam Filtering
Summary. Artificial immune systems (AIS) use the concepts and algorithms inspired by the
theory of how the human immune system works. This document presents the design and initial
evaluation of a new artificial immune system for collaborative spam filtering1 .
Collaborative spam filtering allows for the detection of not-previously-seen spam content,
by exploiting its bulkiness. Our system uses two novel and possibly advantageous techniques
for collaborative spam filtering. The first novelty is local processing of the signatures cre-
ated from the emails prior to deciding whether and which of the generated signatures will
be exchanged with other collaborating antispam systems. This processing exploits both the
email-content profiles of the users and implicit or explicit feedback from the users, and it uses
customized AIS algorithms. The idea is to enable only good quality and effective information
to be exchanged among collaborating antispam systems. The second novelty is the represen-
tation of the email content, based on a sampling of text strings of a predefined length and at
random positions within the emails, and a use of a custom similarity hashing of these strings.
Compared to the existing signature generation methods, the proposed sampling and hashing
are aimed at achieving a better resistance to spam obfuscation (especially text additions) -
which means better detection of spam, and a better precision in learning spam patterns and
distinguishing them well from normal text - which means lowering the false detection of good
emails.
Initial evaluation of the system shows that it achieves promising detection results under
modest collaboration, and that it is rather resistant under the tested obfuscation. In order to
confirm our understanding of why the system performed well under this initial evaluation,
an additional factorial analysis should be done. Also, evaluation under more sophisticated
spammer models is necessary for a more complete assessment of the system abilities.
