<?xml version="1.0" encoding="UTF-8"?>
<mods xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.loc.gov/mods/v3" version="3.1" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-1.xsd">
  <titleInfo>
    <title>Amazon.com new employee access prediction</title>
  </titleInfo>
  <name type="personal">
    <namePart>Thatavarthi, Bhavya Sri</namePart>
    <role>
      <roleTerm authority="marcrelator" type="text">creator</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Chutiporn Anutariya</namePart>
    <role>
      <roleTerm type="text">Chairperson</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Guha, Sumanta</namePart>
    <role>
      <roleTerm type="text">Examination Committee</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Bohez, Erik L.J.</namePart>
    <role>
      <roleTerm type="text">Examination Committee</roleTerm>
    </role>
  </name>
  <name type="personal">
    <namePart>Nicole, Olivier</namePart>
    <role>
      <roleTerm type="text">Exminaiton committee</roleTerm>
    </role>
  </name>
  <name type="corporate">
    <namePart>AIT Fellowship</namePart>
    <role>
      <roleTerm type="text">Scholarship donor</roleTerm>
    </role>
  </name>
  <typeOfResource>text</typeOfResource>
  <originInfo>
    <place>
      <placeTerm type="text">Pathumthani</placeTerm>
    </place>
    <publisher>Asian Institute of Technology</publisher>
    <dateIssued>2019</dateIssued>
    <issuance>monographic</issuance>
  </originInfo>
  <physicalDescription>
    <form authority="marcform">print</form>
    <extent>1 online resource</extent>
  </physicalDescription>
  <abstract>Based on Amazon Inc.'s historical 2010-2011 data, Amazon.com new employee access forecast is based  on  a  system  designed  to replace  resource  administrators on  Amazon.  Our  analysis  shows that  the  given  dataset  with  categorical   values is  very  unbalanced.  Therefore,  during  the preprocessing step, we tried different sampling methods, feature selection, and differentencoding and frequency coding to make the data more suitable for prediction. In the prediction stage we first tested unique models suitable for vector machines with categorical data supportvector machine, logistic regression, light GBM and neural networks. Finally, we combine the best fourprediction results from a random forest, a gradient enhancement and a logistic regression (with encoded data) and an area under curve (AUC) improvement. </abstract>
  <note>A research-study submitted in partial fulfillment of the requirements for the degree of  Master of Engineering in Information Management, School of Engineering and Technology</note>
  <identifier type="uri">http://203.159.5.9/ait-thesis/detail.php?q=B11666</identifier>
  <location>
    <url displayLabel="Full-Text">http://203.159.5.9/ait-thesis/detail.php?q=B11666</url>
  </location>
  <recordInfo>
    <recordCreationDate encoding="marc">      </recordCreationDate>
    <recordChangeDate encoding="iso8601">20260817162202.0</recordChangeDate>
  </recordInfo>
</mods>
