<?xml version="1.0" encoding="UTF-8"?>
<mods xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" version="3.1" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-1.xsd">
  <titleInfo>
    <title>Duplicate record detection for database cleansing</title>
  </titleInfo>
  <name type="personal">
    <namePart>Rehman, Mariam</namePart>
    <role>
      <roleTerm authority="marcrelator" type="text">creator</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Vatcharaporn Esichaikul</namePart>
    <role>
      <roleTerm type="text">Chairperson</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Vilas Wuwongse</namePart>
    <role>
      <roleTerm type="text">Examination Committee</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Jenecek, Paul</namePart>
    <role>
      <roleTerm type="text">Examination committee</roleTerm>
    </role>
  </name>
  <name type="corporate">
    <namePart>Higher Education Commission ( HEC), Pakistan</namePart>
    <role>
      <roleTerm type="text">Scholarship donor</roleTerm>
    </role>
  </name>
  <typeOfResource>text</typeOfResource>
  <genre authority="marc">series</genre>
  <genre authority="marc">technical report</genre>
  <originInfo>
    <place>
      <placeTerm type="code" authority="marccountry">th</placeTerm>
    </place>
    <place>
      <placeTerm type="text">Pathum Thani, Thailand</placeTerm>
    </place>
    <publisher>Asian Institute of Technology</publisher>
    <dateIssued>2009</dateIssued>
    <issuance>continuing</issuance>
  </originInfo>
  <language>
    <languageTerm authority="iso639-2b" type="code">eng</languageTerm>
  </language>
  <physicalDescription>
    <extent>45 p. : ill.</extent>
  </physicalDescription>
  <abstract>Many organizations collect large amounts of data to support their business and decision making processes. The data collected from various sources may have data quality problems in it. These kinds of issues become prominent when various databases are integrated. The integrated databases inherit the data quality problems that were present in the source database. The data in the integrated systems need to be cleaned for proper decision making. Cleansing of data is one of the most crucial steps.  In this research, focus is on one of the major issue of data cleansing i.e. "duplicate record detection" which arises when the data is collected from various sources. As a result of this research study, comparison among standard duplicate detection algorithm, sorted neighborhood algorithm, duplicate elimination sorted neighborhood algorithm, and adaptive duplicate detection algorithm is provided. A prototype is also developed which shows that adaptive duplicate detection algorithm is the optimal solution for the problem of duplicate record detection </abstract>
  <note>A research study submitted in partial fulfillment of the requirements for the degree of  Master of Engineering  Information Management, School of Engineering and Technology</note>
  <note>Research Studies Project Report (M.Eng.) - Asian Institute of Technology, 2009</note>
  <subject authority="lcsh">
    <topic>Database management</topic>
  </subject>
  <relatedItem type="series">
    <titleInfo>
      <title>Research studies project report  ; no. IM-09-01</title>
    </titleInfo>
    <name type="corporate">
      <namePart>Asian Institute of Technology.</namePart>
      <namePart/>
    </name>
  </relatedItem>
  <identifier type="uri">http://203.159.5.9/ait-thesis/detail.php?q=B01412</identifier>
  <location>
    <url displayLabel="Full-Text">http://203.159.5.9/ait-thesis/detail.php?q=B01412</url>
  </location>
  <recordInfo>
    <recordCreationDate encoding="marc">091112</recordCreationDate>
    <recordChangeDate encoding="iso8601">20260818094227.0</recordChangeDate>
  </recordInfo>
</mods>
